Chat Picker

ChatGPT Context Window and Token Limits: What Fits, What Truncates, How to Count Before You Paste

A vendor-sourced explainer on AI token limits: what a token is across ChatGPT, Claude, Gemini and DeepSeek, what counts against a context window, what truncates or degrades, and how to count tokens before you paste.

A token limit is not simply a character limit. Each vendor uses tokens to count input, and the conversion depends on the tokenizer. The practical question is not only “Does the text fit?” but also “What else shares the same window, and which vendor’s tokenizer will count it?”

Chat Picker has not tested, benchmarked or scored any assistant. Everything below is vendor-published, and every figure belongs to the vendor and model that publishes it.

What is a token?

A token is a unit used to count model input. It is not guaranteed to equal one word or one character.

Google’s Gemini token documentation gives a rough English heuristic of about 4 characters per token. The same Google guide says that 60–80 English words occupy about 100 tokens. These are planning approximations, not fixed conversion rules.

DeepSeek publishes different ratios in its API documentation:

  • 1 English character is approximately 0.3 token.
  • 1 Chinese character is approximately 0.6 token.

DeepSeek also says these ratios vary by tokenizer. They should therefore be treated as rough guides rather than exact divisors. Tokenization is tokenizer-specific, so a count produced by one vendor should not be treated as the documented count for another. OpenAI’s tokenizer, for example, is not a universal tokenizer for Claude, Gemini or DeepSeek.

This is why “divide the character count by four” is not a reliable cross-vendor rule. Language, punctuation and the selected tokenizer can change the result.

What counts against a context window?

Anthropic defines the context window as the model’s working memory. It includes the generated response, not just the text supplied by the user, and it is separate from the model’s training corpus.

For Claude, Anthropic documents that the following all consume context:

  • The system prompt.
  • Every message in the conversation.
  • Tool results, images and documents.
  • Tool definitions.
  • The generated output, including extended thinking.

Anthropic’s prompt-caching accounting follows the same window. With caching, input is divided into input_tokens, cache_read_input_tokens and cache_creation_input_tokens. Anthropic says all three categories count toward the context window.

Gemini has a different explicit rule: Google describes its context window as a combined limit for input and output tokens. The input does not have the entire advertised allowance to itself because the generated response also needs room. Google’s token-counting documentation also says tool-definition tokens and system_instruction tokens are included in Gemini’s total_input_tokens.

Vendor-published model figures are model-specific. Anthropic lists certain Claude models with a 1M-token context window, while giving Sonnet 4.5 as a 200K-token example. DeepSeek’s Models & Pricing documentation lists DeepSeek Flash and DeepSeek Pro with 1M context and a 384K maximum output.

Those figures should remain attached to the vendor, model and documentation that publish them. They are not interchangeable limits, and they are not universal across every plan or model. Reusable guidance is to check the target model’s current vendor documentation before relying on a number.

What gets dropped or degraded?

Anthropic says chat interfaces may manage context on a rolling first-in-first-out basis. As conversation turns accumulate, older material can leave the active context while the conversation history itself is preserved completely. In other words, preserving the visible history and fitting the entire history into the model’s working context are different things.

The exact behavior depends on the interface’s context-management policy. A stated context size alone does not say that every interface handles long conversations identically.

There is also a documented quality issue independent of hard truncation. Anthropic states that accuracy and recall can degrade as the number of tokens grows, a phenomenon it calls “context rot.” More context is therefore not automatically better. A larger window can provide more room, but it does not guarantee that a model will use all available material equally well.

The generated side matters too. Anthropic counts Claude’s generated output, including extended thinking, against the window. Gemini likewise treats input and output as a combined allowance. A request that exactly fits on input may still have too little room for the response.

How to count before you paste

Anthropic

Anthropic provides a count_tokens endpoint that returns the total input-token count before the request is sent. Anthropic describes this result as an estimate, which may differ slightly from actual processing.

The target model matters. Anthropic says newer Claude models use a newer tokenizer and can produce approximately 30% more tokens for the same text than earlier Claude models. A count made for an older target should therefore be recalculated for the newer model rather than carried over unchanged.

Google

Gemini’s count_tokens method returns the input-token total and must be called before sending the request. Google says that count includes tool tokens and system-instruction tokens.

Because the result covers input only, it does not include the eventual Gemini output. Gemini’s combined input-and-output context rule still requires room for the response.

OpenAI

OpenAI’s own token-counting documentation uses the tiktoken library. The documented count is:

len(encoding.encode(string))

The encoding must match the target OpenAI model family. OpenAI’s documentation gives encodings such as o200k_base and cl100k_base as examples. Different model families can use different encodings, so the encoding used for one OpenAI model should not be assumed correct for another.

OpenAI also states that it does not endorse third-party tokenizer libraries. Its documented method should not be presented as OpenAI’s endorsement of an unofficial alternative.

DeepSeek

DeepSeek provides an offline demo tokenizer in a downloadable ZIP archive. Its published character ratios are useful for preliminary planning, while the offline tokenizer can be used to estimate a particular input. DeepSeek nevertheless states that the API’s returned usage value is the source of truth.

This distinction is important: an offline count helps answer whether a request appears close to a limit, but it does not replace the usage returned by the API.

Count first, then read the receipt

The reliable workflow is straightforward. Identify the target vendor and exact model. Check that model’s current input and output limits. Count the complete request with the vendor’s documented method, including system instructions, conversation messages, tool material and attached content when applicable. Then leave room for the generated response.

After sending, inspect the response’s usage report. Anthropic and Gemini document usage fields, and DeepSeek explicitly identifies API usage as the source of truth. The local or endpoint count is a planning estimate; the returned usage report is the receipt for what the request actually consumed.

Vendor documentation can establish context sizes, token ratios and counting methods, but it cannot establish a universally best assistant for a particular workload. If the choice matters, run your own trial with representative documents and inspect the resulting usage.

For current prices and limits, see the dated comparison pages.

Compare assistants