Chat Picker

AI

Claude vs ChatGPT vs Gemini: API Cost Comparison

OpenAI and Anthropic publish per-million-token API rates, while Google's cited billing page documents Search-query charges without a comparable Gemini token rate for this comparison.

Sources checked 2 Oct 2026

When you compare Claude vs ChatGPT vs Gemini for API cost, the vendor pages give per-million-token prices for named OpenAI and Anthropic models, while Google’s cited billing details do not support a like-for-like Gemini token-rate row. Chat Picker has not tested the assistants, so this comparison covers published billing terms only.

What the vendors document

Start with API billing rather than a chat subscription. OpenAI’s ChatGPT Plus help page says API usage is separate and independently billed, while Anthropic’s Pro plan help page says Pro does not include Claude Console API usage. The Google AI plans page does not settle whether its subscriptions include API use.

As read in October 2026, the published US-dollar rates include:

Page and model Input per MTok Output per MTok Other documented figure
OpenAI API pricing: gpt-5.6-sol $4.00 $5.00 $0.40 cached input; the row also contains $20.00, $8.00, $10.00, and $30.00 in additional rate columns
Anthropic model overview: Claude Fable 5.1 $10 $50 1M-token context
Anthropic model overview: Claude Opus 5.5 $4 $20 1M-token context
Anthropic model overview: Claude Sonnet 5.5 $2 $10 1M-token context
Anthropic model overview: Claude Haiku 4.5 $1 $5 200K-token context
Google Gemini Developer API pricing Not verified Not verified Billing structure and separate Search-query charges are documented

Keep OpenAI’s additional row values attached to their original column headings rather than treating them as one input/output pair.

Endpoint choices can change the bill. OpenAI’s API pricing page says regional-processing endpoints for models released on or after March 5, 2026, and FedRAMP endpoints carry a 10% uplift. Anthropic’s Claude API pricing documentation says regional and multi-region endpoints carry a 10% premium. It also says cached prompt reads cost less than standard input processing, while Fast mode is a premium research-preview option available through the first-party API for listed models.

Google says a submitted Gemini request may cause one or more Search queries and that each query is charged separately. Its cited page also describes prepaid and pay-as-you-go billing, but the figures verified for this comparison do not form a comparable Gemini token-rate row.

Context capacity is not the same as context economy. Anthropic’s context-window documentation says system prompts, messages, documents, tool definitions, generated output, and extended thinking count toward the window. It also says listed 1M-token models use standard pricing for long-context requests, while accuracy and recall can degrade as token count grows.

Consumer-app limits do not establish API limits. In the same October 2026 reading, OpenAI’s ChatGPT pricing page gives no exact message counts, Anthropic’s Claude pricing page describes rolling five-hour limits plus weekly limits for paid plans, and Google’s Gemini Apps limits page says limits refresh every five hours until the weekly limit is reached.

The consumer-plan pages also make different training statements. OpenAI lists an opt-out for Free, Go, Plus, and Pro. Anthropic lists an opt-out for Free, Pro, and Max and says Team is not trained on by default. Google’s cited plans page does not state a comparable training setting. These are consumer-plan statements, not findings about API retention.

Each vendor also documents accuracy limits. OpenAI’s accuracy note says ChatGPT can be incorrect or misleading and may sound confident when wrong; information beyond its knowledge cutoff requires tools. Anthropic’s Usage Policy calls for reporting potentially inaccurate, biased, or harmful outputs and qualified-professional review for covered advice and recommendations. Google’s Gemini safety guidelines say factually inaccurate output that could cause significant harm should not be generated and that context affects evaluation.

What the documentation cannot tell you

Rate cards do not show which answer will be accurate for your workload, how reliable an assistant will be under load, or what end-to-end latency your users will see. They also cannot predict output length, cache-hit rates, tool calls, Search queries, retries, or the final amount charged for a particular request.

The cited API pages also do not provide a self-hosting cost model. They name no open-weight alternative, license terms, hardware requirement, utilization assumption, or operator-time figure. An open-weight cost advantage therefore cannot be inferred from these hosted API prices.

How to check it yourself

Use the same account region, plan, endpoint, and testing period. Before uploading private material, record the data controls shown in your account and the API terms you accept.

  1. Short structured task. Give each assistant: Classify this message as billing, technical, or account, then return the label and a brief rationale: "I was charged twice after updating my card." Check the classification, unsupported statements, and whether the rationale follows from the message. Write down input, cached-input and output tokens, elapsed time, tool calls, and the charged amount.

  2. Long-document task. Attach a real agreement and prompt: Extract the renewal date, notice period, fee, and cancellation terms. Quote the surrounding text and list anything the document does not state. Look for exact clauses, contradictions, and omissions. Record document size, context use, billed tokens, and cost.

  3. Latency and premium modes. Run the same task under the normal option and any documented fast option. Compare elapsed time, answer structure, token use, retries, and charges. Write down the processing mode used for each run; do not infer latency from a model name.

  4. Search and tool charges. Prompt: Find the current Claude Pro, ChatGPT Plus, and Google AI Pro prices on each vendor's own pricing page. Quote each displayed price and billing period, then link the exact page. Record the date checked, cited pages, Search or tool calls, and any separate charges.

  5. Source fidelity. Attach the same source document to each assistant and use: Using only the attached OpenAI Help Center article, answer: "Can ChatGPT sound confident when it is wrong?" Quote the relevant sentence and name the article. Check quotation accuracy, unsupported additions, and whether the source was actually inspected. Record each discrepancy.

Which rows of the comparison matter

For Claude vs ChatGPT vs Gemini API work, the three-way comparison matrix includes these relevant rows:

  • Input vs. Output Token Pricing: match the exact model, cached-input rate, output rate, and processing column.
  • Context Window Economics: The Hidden Multiplier: separate maximum context from the tokens and output actually billed.
  • Latency-Cost Elasticity: When Speed Costs Premium: include fast modes, regional endpoints, retries, and added tool charges.
  • Provider-Specific Cost Models for Three Scenarios: compare short structured work, long-document work, and tool-assisted retrieval under the same conditions.
  • The Open-Weight Advantage: Self-Hosting vs. API: keep licensing, hardware, utilization, and operator costs in a separate worksheet rather than treating them as zero.

Sources

Current prices and limits for Claude vs ChatGPT vs Gemini, with sources and dates.

Claude vs ChatGPT vs Gemini matrix