Chat Picker

How

Evaluating AI Chatbot Long-Term Memory

ChatGPT, Claude, and Gemini document different memory and past-chat features, but their documentation does not establish real-world recall quality or speed.

Sources checked 2 Oct 2026

Evaluating AI Chatbot Long-Term Memory starts with separating documented controls from actual performance. The OpenAI Memory, Claude memory, and Gemini personalization pages describe different controls but provide no common evidence of comparative accuracy or speed. Chat Picker has not tested the assistants; the previous version’s unsourced statistics and test results have been removed.

What the vendors document

As read on October 2026, OpenAI says Memory can retain relevant details rather than every detail, while Reference chat history can use past conversations. Memory features and controls vary by plan, region, platform, and workspace (Memory in ChatGPT). Data Controls in ChatGPT distinguishes memory from model training. Temporary chats do not create or update memories; turning Memory off does not delete chats; and saved memories are separate. Signed-in ChatGPT Free, Go, Plus, and Pro users can request a data copy, while managed workspaces direct account deletion to the workspace owner. Deleted-memory logs may remain for up to 30 days, and deletion or updates can take time to propagate (Memory in ChatGPT).

Anthropic says Claude can search previous conversations across sessions and cite the originals; past-chat search is available on paid plans. Memory is on by default for Free, Pro, and Max. Team and Enterprise owners control availability, and members enable it. Memory entries survive deletion or expiration of the related conversation, can be deleted individually or reset permanently, and can be exported by admins under organizational retention policies (Claude’s chat search and memory page). Chats and coding sessions are used to improve Claude when users allow this; Incognito chats are excluded (Anthropic’s model-training data page).

Google says Gemini personalization can use past-chat memory, certain connected-app content and activity, and response preferences. It requires a personal Google Account, is unavailable to work, school, and supervised accounts, and is not available to everyone at this time (Get personalization in Gemini Apps). Google’s Gemini Apps Privacy Hub supplements its privacy policy for signed-in processing, but these pages do not settle export, deletion timing, recall criteria, or latency. The Google AI plans page does not state a model-training setting, so Chat Picker marks that row “Not verified.”

For project continuity, OpenAI says shared projects use project-only memory and cannot access a member’s context from outside the project (Projects in ChatGPT). Claude says files placed in a project’s Files section provide persistent reference across conversations (Upload files to Claude). Google says plan upgrades can raise file-upload limits and the number of chats that can reference uploaded files (Upload and analyze files in Gemini Apps).

Published capacity does not prove persistent recall. In pricing pages read on October 2026, OpenAI lists 27K for Instant on Free, 54K for Instant on Go and Plus, and 128K for Pro; reasoning figures are 256K for Go and Plus and 400K for Pro (OpenAI pricing). Anthropic lists up to 1M on every Claude plan, varying by model (Claude pricing). Google gives upload capacity up to 1,500 pages but no app token figure, so Chat Picker marks that cell “Not verified” (Google AI plans).

Accuracy caveats are also explicit. OpenAI says answers can be incorrect or sound confident when wrong, so important information needs checking (OpenAI’s truthfulness page). Anthropic says not to rely on Claude as the sole source and to inspect cited sources (Claude’s accuracy guidance). Gemini sometimes shows related sources (Gemini source guidance); Google says outputs reflect training-data limits and probabilistic variation (Gemini safety guidelines).

For stored personal or project data, OpenAI prohibits unauthorized profiling or distribution of private or sensitive information (OpenAI usage policies). Anthropic requires qualified professionals to review covered advice, recommendations, or subjective decisions directly affecting individuals, along with AI disclosure at the start of consumer chats (Anthropic Usage Policy). Google says Gemini should not produce inaccurate outputs that could cause significant real-world harm (Gemini safety guidelines).

What the documentation cannot tell you

These pages describe controls, not outcomes, and do not define long-term memory with a shared method. Your trial must reveal whether a preference returns without an older value, search finds the correct prior chat, long context is used properly, or edits and deletions propagate. It must also expose stale-memory handling, project handoffs, and export completeness. Use non-sensitive material under the plan and account type you expect to keep.

How to check it yourself

Run the same sequence and record only what you observe:

  1. Context use. Give each assistant this brief: “Create a constraints checklist from the following sections. Goals: prepare a practical neighborhood guide to urban beekeeping. Constraints: prioritize rooftop and balcony options and avoid medical claims. Open decisions: flag anything unresolved. Irrelevant note: the office color is blue. Do not invent answers.” Record retained, omitted, and added constraints, including whether the distractor enters the answer.

  2. Profile persistence. In one session, say: “For future work, remember that I prefer plain English, short decision summaries, and no emoji. Confirm exactly what you can remember.” In a new session ask: “Which preferences did I ask you to remember earlier? Quote them and identify the source.” Record exact, inferred, and missing items.

  3. Recall latency. Ask: “Retrieve the beekeeping guide’s goal and constraints from our previous conversation, then identify the source.” Time the interval from sending the prompt to receiving the response. Record elapsed seconds, correctness, and whether a source link or citation appears.

  4. Portability and export. Ask: “List every stored preference you can access. Distinguish explicit items from inferences and show where I can inspect, export, edit, or delete them.” Inspect the interface and record the export format, included items, omissions, and errors.

  5. Editing and deletion. Ask: “For this test, remember: use green for status labels.” If a control appears, edit or delete the entry. In a new session ask: “Is ‘use green for status labels’ currently saved? Cite it or say no.” Record what the interface and new response show, including delays or stale text.

  6. Project continuity. Create a handoff with: “Our community-garden documentation survey is complete, the planting schedule is approved, and budget questions remain open. Summarize decisions and open tasks without inventing owners.” Later request a fresh handoff. Record omissions, contradictions, invented tasks, and whether project context carries across sessions.

Which rows of the comparison matter

On the Claude vs ChatGPT vs Gemini matrix, read “Context window in the app” beside “How usage limits are described.” Check “Free plan,” “Main paid plan,” “Heavy-use plans,” and “Team plan” for feature availability, then use “Training on your chats” to separate personalization from model training. Upload or storage figures do not establish cross-session recall, and “Not verified” cells should remain unresolved. Recheck vendor pages before subscribing.

Sources

For current prices and limits, see the dated comparison pages.

Compare assistants