How
AI Chat Assistants for Social Survey Design and Analysis
A practical guide to vendor-documented limits for ChatGPT, Claude and Gemini, plus a self-check for social survey work that does not assume comparative answer quality.
Sources checked 3 Oct 2026
ChatGPT, Claude and Gemini vendor pages settle some practical questions about social survey design and analysis: supported files, analysis features, prices, limits and data handling. They do not establish comparative answer quality, and Chat Picker has not tested the assistants for this use. Use the documentation comparison and self-check below; it is not a ranking.
What the vendors document
The vendor figures below are stated as read on October 1, 2026.
- ChatGPT. OpenAI's ChatGPT pricing page lists Plus at $20 per month, a training opt-out on Free, Go, Plus and Pro, 27K of Instant app context on Free, and 400K of reasoning context on Pro. OpenAI's File Uploads FAQ caps each chat file at 512MB, a CSV or spreadsheet at about 50MB, and Free usage at three uploads per day. OpenAI's data analysis page says ChatGPT can run Python in a stateful Jupyter environment and make charts, but it cannot make external web or API calls there. OpenAI's accuracy guidance says answers can sound confident while wrong and important quotes, data and references need checking.
- Claude. Anthropic's Claude pricing page lists Pro at $20 per month, up to 1M context on every plan depending on model, and a training opt-out on Free, Pro and Max; Team is not trained on by default. Claude's file upload page gives 500MB per chat file and 30MB per project file; XLSX uploads require code execution and file creation. Claude's incorrect-response guidance says Claude can produce convincing but ungrounded quotations and should not be the only source of truth.
- Gemini. Google's U.S. AI plans page lists a free plan with a Google Account and 15 GB storage; Google AI Pro is $19.99 per month with 4x the Free usage limits and 5 TB storage. The same page gives no app token figure and does not state a training setting. Gemini's file upload page says a prompt can contain up to 10 supported files, each non-video file can be 100 MB, each video can be 2 GB, and the web app can make customizable charts from uploaded spreadsheets. Gemini's safety and policy guidelines acknowledge that outputs may reflect limited viewpoints or overgeneralizations and that LLM outputs are probabilistic.
- Sensitive data. OpenAI's Usage Policies prohibit unauthorized aggregation, monitoring, profiling or distribution of private or sensitive information. Anthropic's Usage Policy says elevated-risk uses should integrate relevant human expertise and that covered advice, recommendations and subjective decisions require qualified professional review before finalization. It also requires AI involvement to be disclosed at the beginning of a session when output is presented directly to individuals or consumers. These are vendor rules, not a determination that a survey is lawful or ethical.
What the documentation cannot tell you
The vendor pages do not report whether a construct fits your theory, a sample frame covers the intended population, questionnaire wording is neutral, generated code preserves the data, or open-text codes retain context. Context and file ceilings also do not show how your actual missingness, languages, row count or upload mix will fit a session.
OpenAI's pricing page and Anthropic's pricing page publish no exact message counts, while Google's AI plans page describes compute-based limits that refresh every five hours up to a weekly limit. Test candidates with identical synthetic or de-identified materials and prompts, and keep a shared log of edits, code runs, source checks and open questions.
How to check it yourself
If comparing assistants, use identical source files and prompt wording, record each account and plan, and give every assistant the same revision opportunity. Keep one shared audit log.
- Define the construct. Give: “Define the construct ‘trust in local public libraries among adult users.’ State its boundary, dimensions, observable indicators, neighboring constructs to exclude and assumptions needing human review.” Look for a measurable definition, not a bundle of causes. Write down the retained definition, indicators, exclusions and unresolved assumptions.
- Audit the sampling frame. Give: “Using the attached de-identified roster, audit it as a sampling frame. Report duplicate IDs, missing eligibility fields, out-of-scope records and possible coverage gaps. Do not call it representative; list what population information is needed to judge coverage.” Look for separation of record defects from coverage questions. Write down cleaning rules, issue counts, exclusions and missing source documents.
- Refine questionnaire items. Give: “Review the attached pilot items about library trust. Preserve the stated construct and response scale. Identify double-barreled, leading, overlapping, missing and unbalanced response options. Give a reason and revised wording for every change, and add skip logic where needed.” Look for preserved measurement meaning and explainable edits. Record each proposed change, its reason and every point requiring a researcher decision.
- Check pilot-data code. Give: “Using the attached de-identified pilot.csv and data dictionary, write Python to check duplicate IDs, missing values, values outside documented ranges and inconsistent field types. State assumptions, report exclusions and do not silently drop rows.” Look for explicit validation, reproducible logic and separation between data checks and interpretation. Save the code, assumptions, warnings and results you independently rerun.
- Code open-ended responses. Give: “Using the attached de-identified CSV with open_text, human_code and reviewer_note fields, draft a codebook, apply it and list ambiguous responses. Preserve the original text and do not infer protected characteristics.” Look for traceable labels, overlapping themes and explicit handling of disagreement. Write down the codebook version, exclusions, disagreements and original responses checked.
- Review ethics and privacy. Give: “Review the attached data dictionary, consent draft and proposed community-survey workflow. Flag direct identifiers, sensitive fields, linkability, access roles, retention choices, reidentification risks and mismatches between consent promises and data use. End with items requiring human, privacy or legal review.” Look for data minimization and clear ownership of each control. Record the issue, proposed action, owner and approval status; treat the response as a screening list, not approval.
Which rows of the comparison matter
Open the Claude vs ChatGPT vs Gemini matrix. For survey work, read the Free plan, Low-cost tier, Main paid plan, Heavy-use plans, Team plan, Context window in the app, How usage limits are described, and Training on your chats rows.
The matrix was checked on October 2, 2026, and its figures carry vendor sources and read dates. A cell marked “Not verified” is unconfirmed, not evidence that a feature is unavailable. Use the Chat Picker method page to understand the sourcing and confirm current vendor details before subscribing.