Chat Picker

How

Testing AI Chatbot Knowledge Breadth

Explains the accuracy warnings, tool limits, and data controls documented for ChatGPT, Claude, and Gemini, then provides a repeatable 30-question knowledge-breadth test.

Sources checked 2 Oct 2026

Testing AI chatbot knowledge breadth means giving assistants comparable cross-disciplinary questions and checking the answers against dependable sources. The vendors’ own pages document feature access, limits, data controls, and accuracy warnings, but they do not establish a head-to-head breadth result. Chat Picker has not tested ChatGPT, Claude, or Gemini for this purpose, so this page gives you a controlled test instead of a score, ranking, or winner.

What the vendors document

The earlier version’s unsourced statistics and test results have been removed. The feature and limit statements in this section are as read on October 1, 2026.

Search and research

OpenAI says web search is available across ChatGPT Free, Go, Plus, Pro, Business, Enterprise, and Edu, as well as for signed-out users, subject to plan limits in OpenAI’s web search help page. Its Deep Research help page says the feature can use uploaded files, the public web, selected sites, and enabled apps. Completed research outputs include citations or source links, while usage varies by plan.

Anthropic says Research is available to paid Claude Pro, Max, Team, and Enterprise users. It requires web search, can use connected context such as Gmail, Google Calendar, and Google Docs, and may consume standard usage limits faster because it retrieves multiple sources. See Anthropic’s Research help page.

Google says Gemini Deep Research requires sign-in and an age of 18 or older. Google Search is included by default, and users can add sources such as personal Gmail, Drive, uploaded files, or NotebookLM notebooks. The Google Deep Research help page does not state a plan requirement beyond those conditions. These access paths do not establish how much any assistant knows.

Files and context

For app context, OpenAI’s pricing page lists 27K on Free, 54K on Go and Plus, and 128K on Pro for Instant models. Its listed Reasoning-model figures are 256K on Go and Plus and 400K on Pro. Anthropic’s pricing page says Claude can offer up to 1M context on every plan, depending on the model. The Google AI plans page does not publish an app token figure, so Chat Picker marks that cell “Not verified.”

File support also differs. OpenAI’s File Uploads FAQ says uploads are available on Free and paid plans, subject to account and plan limits. Anthropic’s file upload help page allows files in individual chats or persistent projects. Google’s file help page says signed-in Gemini users can upload documents, spreadsheets, notebooks, photos, and videos, subject to availability and account settings.

Data, accuracy, and policy

The training-control rows are not the same. OpenAI’s pricing page says an opt-out is available for Free, Go, Plus, and Pro chats. Anthropic’s pricing page says an opt-out is available on Free, Pro, and Max, while Team is not trained on by default. Google’s plans page links to data-handling information but states no corresponding setting, so Chat Picker does not mark that control as verified.

All three vendors warn about incorrect output. OpenAI’s accuracy note says ChatGPT may be wrong and that confidence is not reliability; its search documentation also warns that citations can be incomplete, outdated, or incorrect. Anthropic’s accuracy note says Claude should not be the only source of truth and that original sources may add missing context. Google’s safety guidelines say outputs can reflect training-data limits, limited viewpoints, and overgeneralizations.

Usage policies also matter for consequential questions. OpenAI’s usage policies do not replace professional duties and restrict tailored licensed advice without appropriate professional involvement. Anthropic’s Usage Policy requires qualified professional review for covered advice, recommendations, and decisions before dissemination or finalization, along with AI disclosure in specified consumer-facing uses. These rules do not quantify answer accuracy.

What the documentation cannot tell you

The pages do not reveal whether an assistant will connect ideas across fields, interpret an unfamiliar source correctly, distinguish evidence from inference, or handle your particular combination of files and tools. Source quality, plan access, prompt wording, and account settings can all change what you receive.

Your trial can show behavior under your conditions, but it cannot establish a universal knowledge ranking. Use it to assess suitability for research, study, source review, or drafting while keeping the subject, tools, and files consistent.

How to check it yourself

  1. Create and lock the questions. Build a 30-question set spanning engineering, science, public policy, history, economics, and software. Give every assistant the same brief: “Answer these questions from your existing knowledge, state uncertainty, and do not invent sources.” Look for consistent scope. Record the exact wording, topic, source checklist, plan, and account state.

  2. Run a closed-book round. Give each prompt in a fresh chat with search, research, connected apps, and uploads disabled where the controls allow it. Look for correct terminology, stated assumptions, cross-domain connections, contradictions, and unsupported specifics. Record each response as verified, partly supported, unsupported, or unclear.

  3. Repeat with research tools. Use the same prompts in new chats with each vendor’s native search or research function. Look for claim-level citations, source dates, primary material, conflicting evidence, and explicit uncertainty. Open the cited pages and record whether each source supports the associated claim rather than merely sharing its keywords.

  4. Test file grounding. Give each assistant the same report and ask: “Use the attached report to answer the questions below. Cite page numbers, distinguish direct evidence from inference, and state anything the report does not establish.” Look for correct handling of charts, tables, qualifiers, and contradictions. Record the supporting passages for every important conclusion.

  5. Audit and interpret. Prepare authoritative excerpts, then send: “Compare your answer with the excerpts below. Correct every unsupported statement and explain any disagreement.” Look for acknowledged errors and evidence-based disagreement. Record corrections by field and test mode; do not convert category counts into a general knowledge score.

Which rows of the comparison matter

In the Claude vs ChatGPT comparison matrix, start with Free plan, Low-cost tier, Main paid plan, Heavy-use plans, and Team plan. Then check Context window in the app, How usage limits are described, and Training on your chats.

Those rows show whether your intended plan has the documented capacity and controls for sustained testing. Confirm current prices and limits on the linked vendor pages. Treat a blank “Not verified” cell as unknown, not as zero, unavailable, or absent.

Sources

For current prices and limits, see the dated comparison pages.

Compare assistants