Chat Picker

How

How to Train or Fine-Tune an AI Chat Model

It contrasts documented weight tuning with saved customization and gives readers a six-part evaluation plan they can run themselves.

Sources checked 2 Oct 2026

Fine-tuning changes a model’s behavior with task-specific examples, while custom instructions, memory, and Gems customize an assistant without showing that its weights were retrained. For this use, OpenAI’s supervised fine-tuning guide says its platform is winding down and closed to new users; Google’s Gemini tuning page says no available model supports fine-tuning in the Gemini API or AI Studio; and the Anthropic pages read here document memory and preferences but do not say that a weight-tuning workflow is available. Chat Picker has not tested ChatGPT, Claude, or Gemini for model training.

What the vendors document

  • ChatGPT and OpenAI. As read on October 1, 2026, OpenAI says the fine-tuning platform is winding down, is inaccessible to new users, and lets existing users create training jobs for the coming months. Training files use JSONL, the minimum is 10 training examples, and OpenAI recommends starting with 50 well-crafted demonstrations. It also recommends setting up evaluations and holding out data before training. Fine-tuned models remain available for inference until their base models are deprecated, while the Responses API stores model responses for 30 days by default. OpenAI’s data-control documentation allows eligible signed-in consumer users to turn off model improvement, although doing so does not delete saved chats. It also says personalized Temporary Chats can use existing memories and custom instructions, subject to account and workspace settings. OpenAI’s accuracy note says answers can be incorrect or misleading and important information should be checked against reliable sources.

  • Claude and Anthropic. Claude’s memory documentation says memory defaults on for Free, Pro, and Max, while Team and Enterprise owners control availability. Anthropic’s model-training privacy notice says chats and coding sessions are used to improve models when the user allows it. Data can include the related conversation, custom styles, and conversation preferences, while Incognito chats are excluded even when model improvement is enabled. Anthropic’s accuracy note says Claude can produce incorrect or misleading responses, advises against using it as the only source of truth, and says cited sources should be reviewed in context.

  • Gemini and Google. Google says no available model supports fine-tuning in the Gemini API or AI Studio. Support exists in Gemini Enterprise Agent Platform, but Google states that it has no immediate plans to add it elsewhere. A Gemini Gem is a named set of instructions with optional files. Using Gems requires signing in, while Gemini personalization requires a personal Google Account. Gems are subject to Google’s Terms of Service and Prohibited Use Policy, and Google advises discretion before relying on or publishing generated content. Google’s source-display documentation explains when a Sources button appears, but it does not provide an accuracy warning. The Gemini Apps privacy material read here does not state a model-training control.

For an open-weight route, Hugging Face’s PEFT documentation describes adapting models by tuning a small number of added parameters rather than all parameters. llama.cpp documents a plain C/C++ implementation, an OpenAI-compatible server, and several integer quantization options. Ollama’s FAQ says its local operation does not see prompts or data, while cloud-hosted models process prompts and responses to provide the service without storing, logging, or training on their content. These are documented mechanisms, not evidence that any route will meet your quality or hardware target.

What the documentation cannot tell you

Vendor pages describe features, controls, and limits, but they cannot establish whether your examples represent real use, whether a model retains the intended behavior, or whether a setup fits your hardware and budget. OpenAI’s model optimization guide also says LLM output is non-deterministic and behavior can change between model snapshots and families, making your own repeated evaluation necessary.

Treat customization and fine-tuning as separate baselines. A custom instruction, saved preference, memory entry, or Gem may address a workflow, while a documented fine-tuning job updates model weights. Your trial must still expose unsupported assumptions, formatting failures, latency, memory use, and cost.

How to check it yourself

  1. Choose the base and adaptation path. Start with this test prompt: “Classify the support ticket ‘I was charged twice and cannot log in’ as billing, access, or unknown. Return only JSON with intent and a one-sentence reason.” Give the same fixed cases to each base model with a plain prompt, then with custom instructions, saved memory, or a Gem where available. Look for wrong labels, invalid JSON, and invented facts. Record each error and any change caused by customization.

  2. Prepare the training dataset. Give the training job paired inputs and target outputs while reserving a separate holdout set. Include a stable row ID with each pair, then ask an assistant to review the pairs against the labeling rules. Look for whether its review cites the correct rows and catches duplicate inputs, conflicting targets, missing fields, formatting errors, and exposed secrets. Record every accepted or rejected flag and the reason for excluding a row.

  3. Set up the environment. Give the assistant your exact operating system, installed RAM, GPU, context target, expected concurrent requests, and network requirement. Ask for a setup plan, then verify each step against the tool documentation. Look for unsupported hardware claims, unclear model licenses, and hidden cloud dependencies. Record software versions, model source, context setting, and whether prompts remain local.

  4. Execute the fine-tuning run. Give the training pipeline one fixed configuration and change one variable at a time. Save the configuration, logs, checkpoints, errors, duration, and memory measurements. Look for whether each artifact loads and exposes the expected interface. Record every deviation from the planned setup rather than silently correcting it.

  5. Evaluate the result. Give the base, customized, and fine-tuned versions identical holdout prompts, expected labels, and a written scoring rubric. Look for correct classifications, valid structure, omissions, unsupported claims, and unnecessary refusals. Record exact-match results, formatting failures, latency, and any API charges shown in your logs.

  6. Test deployment and inference. Send the same requests through your intended serving stack. Change one setting at a time, such as context length or quantization. Look for output changes, cold-start delays, throughput limits, peak memory, request failures, and rollback problems. Record each setting beside its measurements so you can reproduce the deployment later.

Which rows of the comparison matter

Open the Claude vs ChatGPT vs Gemini matrix. For this use, focus on Free plan, Main paid plan, Heavy-use plans, Context window in the app, How usage limits are described, Training on your chats, and Team plan. These rows help with access, cost, context, usage, chat-data controls, and workspace settings.

They do not settle whether fine-tuning will work for your task. An app’s context window also does not establish an API training limit, and the “Training on your chats” setting is separate from uploading a training dataset. Recheck vendor pages before subscribing or sending sensitive material.

Sources

For current prices and limits, see the dated comparison pages.

Compare assistants