AI
AI Assistant Personalization: Fine-Tuning Compared
A documented comparison separates model-weight updates from saved instructions, reusable Gems, and memory, showing that availability and privacy controls differ across ChatGPT, Claude, and Gemini.
Sources checked 2 Oct 2026
AI assistant personalization is not one feature: vendor pages separate saved instructions, reusable Gems, memory, and model-weight updates, and they do not make those tools interchangeable. OpenAI says its API fine-tuning platform is winding down, Google says no available Gemini API or AI Studio model supports fine-tuning, and the Anthropic pages read here do not document a consumer Claude fine-tuning route. Chat Picker has not tested ChatGPT, Claude, or Gemini for personalization; this page compares their documented availability, controls, and limits.
What the vendors document
OpenAI’s Supervised fine-tuning guide, as read on October 1, 2026, says the fine-tuning platform is winding down and closed to new users, while existing users can create training jobs for the coming months. It sets 10 as the minimum number of training examples, recommends starting with 50 well-crafted demonstrations, and advises creating evaluations and setting aside holdout data before training. OpenAI says the appropriate number depends on the use case.
Google’s Gemini API fine-tuning guide, read on October 1, 2026, says no available model supports fine-tuning in the Gemini API or AI Studio. It says fine-tuning is supported in Gemini Enterprise Agent Platform, but there are no immediate plans to add support elsewhere. A different feature is documented in Google’s Gems guide: a Gem has a name and instructions and can include files. The same page says Gems on personal Google Accounts are scheduled to transition to skills starting in November 2026.
Consumer personalization is separate from API fine-tuning. OpenAI’s Memory guide says ChatGPT can use relevant preferences and details from chats when Memory is enabled, but it does not retain every detail from every conversation. Its Data Controls guide says a personalized Temporary Chat can use existing memories and custom instructions when available; Temporary Chats do not create memories or improve models.
Anthropic’s Claude memory guide documents memory and past-chat search rather than a consumer model fine-tuning method. It says memory defaults on for Free, Pro, and Max, while Team and Enterprise owners control availability. Anthropic’s model-training privacy article says consumer chats and coding sessions are used to improve models when the user allows it, while Incognito chats are not used even when Model Improvement is enabled.
Google’s Gemini personalization guide says personalization can use past chats, connected Google app activity, response preferences, and instructions. It requires a personal Google Account and says the features are unavailable to work, school, and supervised accounts, as well as not being available to everyone. The vendor pages read here do not settle whether Gemini app chats have an app-level model-training opt-out.
As read on October 1, 2026, OpenAI’s Plus page lists Plus at $20 per month billed monthly, Anthropic’s pricing page lists Pro at $20 per month or $17 per month with annual billing at $200 up front, and Google’s US plans page lists AI Pro at $19.99 per month. Confirm the current price and included features before subscribing.
For open-weight models, the tooling is separate from a hosted assistant. Hugging Face’s PEFT documentation says PEFT adapts pretrained models by tuning a small set of additional parameters rather than all parameters, reducing compute and storage requirements. The ChatGLM3 documentation says it provides a basic framework for fine-tuning ChatGLM3-6B with a user’s dataset. Ollama’s FAQ says local operation does not see prompts or data, while its cloud service processes prompts and responses without storing, logging, or training on their content.
The accuracy warnings also differ. OpenAI says ChatGPT can sound confident while being wrong and recommends checking important information. Anthropic says users should not rely on Claude as their only source of truth and should inspect cited sources. Google advises discretion when relying on generated Gems and says Gems are subject to Google’s Terms of Service and Prohibited Use Policy.
What the documentation cannot tell you
The cited pages do not compare how consistently the three assistants apply a saved preference, recall a detail after a gap, or handle a specialized workflow under the same conditions. Memory documentation describes controls and availability, not a guarantee that a particular detail will be remembered or used at the right time.
Your trial must also reveal what your plan, region, device, and workspace expose. Privacy settings show what the vendor says it can do; only your own records show which controls you can find and how deletion behaves. Open-weight setups add model, hardware, quantization, and serving choices that are not settled by a hosted-plan comparison. The older version of this article contained unsourced statistics and test results; those have been removed.
How to check it yourself
-
Test explicit instructions. Where custom instructions or Gems are available, enter: “For the next five requests, call me Alex, explain unfamiliar technical terms, put a three-bullet decision summary before the details, and ask one clarifying question when my request has two plausible interpretations.” Then ask, “Explain why a database index can speed up a query.” Record the exact setting, wording, and whether the rules appear in a new chat.
-
Test cross-session memory. First send: “For this test, remember that my fictional client Apex uses a Tuesday launch review and prefers decisions before background. Do not infer other preferences.” In a new chat, ask: “What launch-review day and response order did I give for the fictional client Apex?” Write down the answer, any invented detail, and any source shown.
-
Test a specific use case. Give this prompt: “Turn these meeting notes into a decision brief: ‘The beta opens Friday; support is ready; legal approval is pending; target date is August 14.’ Include the decision, owner, blocker, next action, and unresolved question.” Look for adherence to each requested element, then repeat it in a fresh chat and record any changes.
-
Check the privacy trade-off. Do this with non-sensitive test data. Send: “For this test account, remember that the fictional project code is Cedar.” Review the available memory and model-improvement controls, then test whether you can delete the entry. Record what appears, what deletion confirms, and which settings governed the test.
-
Evaluate a tuning route. Use a fixed set of representative training examples and separate holdout prompts. One training example could be: “Ticket: ‘My refund is missing.’ Return the support intent and priority.” Then test: “Ticket: ‘The app closes when I export.’ Return the support intent and priority.” Record the data, configuration, failures, cost, and date; do not treat one successful answer as proof of reliability.
Which rows of the comparison matter
Open the ChatGPT vs. Claude vs. Gemini matrix and focus on these rows:
- Free plan: whether the personalization feature is available before payment.
- Main paid plan: current price and documented plan-level access.
- Context window in the app: whether a preference test needs large files or long conversations.
- How usage limits are described: whether repeated testing may encounter tool or message limits.
- Training on your chats: the documented choice or default affecting model improvement.
Chat Picker leaves unconfirmed figures blank and marks them “Not verified,” rather than filling gaps. Check the sourcing method and the linked vendor pages before making a subscription decision.