How
AI Tools for Education: Teaching and Assessment
Explains what vendors document for AI-assisted teaching and assessment, including features, usage limits, data terms, and the need to verify outputs.
Sources checked 2 Oct 2026
AI Tools for Education: Teaching and Assessment should be compared against your curriculum, assessment, privacy, and budget requirements, not an assumption that one assistant is better. OpenAI’s truthfulness guidance says ChatGPT may sound confident when wrong; Anthropic’s guidance on Claude responses says Claude can display convincing quotations that are not grounded in fact; and Google’s Gemini policy guidelines say its outputs can reflect limits in training data. Chat Picker has not tested the assistants for education or ranked them.
What the vendors document
The plan figures below, as read on the linked vendor pages in October 2026, cover these documented differences.
-
ChatGPT: OpenAI’s Free tier FAQ says Free includes unlimited everyday text chats, subject to abuse-prevention safeguards, while uploads, images, voice, and data analysis have separate limits. OpenAI’s ChatGPT Plus page lists Plus at $20 per month, with file uploads and analysis, broader model options, and advanced reasoning. Message caps can vary with system conditions, and API usage is separate and billed independently.
-
Claude: Anthropic’s Pro plan article lists Pro at $20 per month in the US. Claude’s pricing page says Free includes web search, file creation, and code execution, while Pro adds Docs, Slides, projects, and more models. Usage limits reset on a rolling five-hour window, and paid plans also have weekly limits.
-
Gemini: Google’s AI plans page lists Free at $0 per month with 15 GB of cloud storage across Gmail, Drive, and Photos. Google AI Pro is listed at $19.99 per month with 4x Free usage access and 5 TB of storage. Gemini Apps limits help says limits depend on prompt complexity, features, and chat length, refresh every five hours until the weekly limit is reached, and can be extended with AI credits.
-
Accuracy and sources: OpenAI says ChatGPT can produce incorrect or misleading information and recommends using it as a first draft rather than a final source. Anthropic says Claude may produce authoritative-sounding but incorrect statements; when using web results, check the original pages because Claude’s synthesis may omit context. Google’s related-sources guidance says Gemini Apps sometimes show sources within or below a response, including public websites, uploaded files, and connected Workspace documents or email. A source panel provides material to inspect; it does not prove the conclusion.
-
Data: OpenAI’s Business overview says OpenAI does not train on Business workspace data and directs needs such as BAAs or Zero Data Retention to a contracted offering. Anthropic’s pricing page lists model training as opt-out for Free, Pro, and Max, while Team is not trained on by default. Google’s plans page does not state a chat-training setting, so Chat Picker’s comparison matrix marks that item as Not verified.
-
Use policies: OpenAI’s usage policies prohibit automated high-stakes decisions in sensitive areas without human review. Anthropic’s Usage Policy requires qualified professional review for covered advice and subjective decisions that directly affect people, plus AI disclosure in specified consumer-facing sessions. Google says context, including educational applications, matters, while its probabilistic models can produce new and different responses.
What the documentation cannot tell you
The vendor pages do not establish standards alignment, assessment validity or bias resistance, rubric consistency, useful feedback, latency, or compatibility with named LMS platforms. They also do not determine whether a plan meets your institution’s age, accessibility, retention, security, or recordkeeping rules.
A source panel can help you inspect material, but it does not validate a mark or claim. Those questions require a controlled pilot and institutional review. The previous version’s unsourced statistics and test results have been removed.
How to check it yourself
-
Curriculum alignment. Give the assistant a complete lesson plan and the exact standards text. Ask: “Map each objective to the supplied standard, quote the words that support the match, and label any unsupported match.” Look for invented requirements, weak matches, and missed standards; write down every correction you would make.
-
Assessment validity and bias. Give it de-identified sample responses and a fixed rubric. Ask: “Score each response only against this rubric, cite supporting words, and do not infer personal characteristics.” Look for criterion-level reasoning and unsupported inferences; record every score you cannot defend from the response.
-
Feedback and rubric consistency. Give the same response and rubric in separate fresh chats. Ask: “Return a criterion-by-criterion score, one improvement, and one preserved strength, with evidence from the response.” Look for reasons behind each score; time each run yourself and record elapsed time, score changes, and feedback that shifts without a textual reason.
-
Privacy and controls. Use made-up learner data, not real student records. Ask: “List the inputs you need, the output you will return, and the settings I should verify before entering student information.” Look for assumptions and missing cautions, then verify the plan and administrative controls yourself. Record training settings, retention and deletion terms, export options, and required agreements.
-
LMS workflow. Give the assistant a made-up gradebook export. Ask: “Convert these rows to a CSV that can be imported again, preserve the headers, and flag uncertain values.” Look for destructive changes and uncertain formatting. Test the file in a nonproduction LMS environment and record connector availability, manual steps, errors, and data loss.
-
Cost and scale. Give the same representative lesson, rubric, and response during a realistic teaching week. Look for tool restrictions, reached caps, and extra billing prompts. Record subscription seats, API charges where applicable, staff correction time, and work left unfinished.
Which rows of the comparison matter
On the Claude vs ChatGPT vs Gemini matrix, read the rows for Free plan, Low-cost tier, Main paid plan, Heavy-use plans, Team plan, Context window in the app, How usage limits are described, Training on your chats, and Ads. Also check file handling, storage, source display, and API billing where the matrix provides them.
Treat a blank or “Not verified” cell as unanswered, not as zero or a negative answer. Context size is capacity, not evidence of assessment quality. Confirm current prices, limits, and institutional requirements on the vendor page before subscribing.