Top
Open-Source ChatGPT Alternatives Compared for Self-Hosting
A documentation-based comparison separates hosted assistant plans from open-weight model access and self-hosting requirements, without claims about untested answer quality.
Sources checked 3 Oct 2026
If you are choosing the best AI chatbot to operate on your own systems, OpenAI’s pricing page, Anthropic’s pricing page, and Google’s US AI plans page, as read on October 1, 2026, settle hosted plan terms—not self-hosting rights for ChatGPT, Claude, or Gemini. Separate project pages document open-weight models and local deployment tools. Chat Picker has not tested them for this use, so this is a documentation comparison, not a quality ranking.
What the vendors document: best AI chatbot limits
Hosted plans differ. OpenAI’s pricing page lists limited uploads, images, voice and deep research on Free, plus ChatGPT Plus at $20 per month with monthly billing (read October 1, 2026). Anthropic’s pricing page lists Claude Pro at $20 monthly or $17 monthly with $200 paid upfront for annual billing; it lists no paid tier between Free and Pro (read October 1, 2026). Google’s US plans page lists AI Plus at $4.99 per month, AI Pro at $19.99, and AI Ultra at $99.99 or $199.99; it says compute-based limits refresh every five hours, subject to a weekly limit (read October 1, 2026). None of these cited pricing pages says that the corresponding ChatGPT, Claude or Gemini model can be downloaded and self-hosted.
Data choices also remain tied to those services. OpenAI says a training opt-out is available on Free, Go, Plus and Pro. Anthropic says an opt-out is available on Free, Pro and Max, while Team is not trained on by default. Google’s plans page states no corresponding training setting. These are documented service settings, not findings about a model running on your hardware.
The accuracy guidance is similarly product-specific. OpenAI’s accuracy note says ChatGPT can produce incorrect or misleading answers and may sound confident when wrong, so important information should be checked against reliable sources. Anthropic’s accuracy note warns that Claude can make convincing but incorrect statements and should not be the only source for high-stakes decisions. Google’s source guide explains where a Sources button and related links appear when available; it does not guarantee that a response is correct.
The projects read for this page are not the whole field. Among the model projects, Meta’s Llama models repository lists Llama 3.1 70B with a 128K context length (read October 1, 2026). It requires license acceptance and an access request, and says the weights are licensed for researchers and commercial entities. The page does not give a Llama 3.1 70B memory minimum. Mistral Docs identifies Mistral Large 3 as an open-weight, general-purpose multimodal model, but the model and pricing pages do not set out its exact self-hosting license, hardware floor or installation steps (both read October 1, 2026). The ChatGLM3 repository says its weights are open for academic research and allow free commercial use after registration. Its shown default FP16 load uses about 13GB of VRAM; CPU inference uses about 32GB of memory and is slower. It also says four-bit quantization causes some performance loss (read October 1, 2026).
The deployment tools have different jobs. Ollama’s repository documents a REST API and integrations for local and cloud models. Its FAQ says a local Ollama instance does not see prompts or data, while prompts sent to cloud-hosted models are processed to provide the service but are not stored, logged or used for training; basic account information and limited metadata are still collected (read October 1, 2026). Open WebUI describes a self-hosted platform that can run offline with Ollama or OpenAI-compatible APIs. Its :main image includes retrieval, document and voice components, while :slim expects external services for features such as knowledge search and voice.
LM Studio supports macOS, Windows and Linux, can operate offline after model files are obtained, and supports local models through llama.cpp and Apple’s MLX. The llama.cpp repository documents prebuilt binaries, Docker and source builds, along with quantization and CPU-plus-GPU hybrid inference. PEFT is an adaptation library rather than a chatbot runtime: it updates a small set of extra parameters instead of fine-tuning every parameter.
What the documentation cannot tell you
Project documentation does not reveal how a model will handle your particular files, code, terminology or failure cases. It also cannot establish your actual memory headroom, latency, throughput, tool reliability or compatibility with a particular setup.
Vendor accuracy warnings do not automatically transfer from a hosted assistant to every local model. Your own trial can document behavior under defined conditions, but it cannot prove universal reliability. Chat Picker has no test results or scores to fill that gap.
How to check it yourself
Before testing, save the exact model revision, license acceptance, quantization and backend. If the project page does not specify a license or hardware requirement, mark it unknown rather than inferring it from the model name.
-
Test grounded extraction. Give the assistant: “Read this test note: The west gate opens only for Mara. The silver key lies beneath the elm. The red ledger stays in the archive. Return JSON with the gate rule, key location and ledger location, then quote the supporting sentence for each value.” Check exact extraction, valid structure and whether quotations match the text. Record omissions, additions and formatting failures.
-
Test code generation. Ask: “Write a Python function named
title_case_wordsthat accepts a list of strings, preserves punctuation and empty input, and returns every word capitalized. Include assertions for ordinary input, punctuation and an empty list.” Run the code and inspect its behavior. Record which assertions pass, any exceptions and any unnecessary changes to the requested function. -
Test document retrieval. Attach a document you have permission to use, then ask: “Answer only from the attached document: What is its purpose, who owns the process, and what approval is required? For each answer, give an exact supporting passage and page number. Write ‘Not stated’ if an answer is absent.” Record unsupported additions, weak quotations and incorrect page references.
-
Test structured tool use. Give this prompt: “For the sentence ‘Mara opens the west gate before sunrise,’ return one JSON tool call named
request_gate_accesswith fields for user, gate and requested time. Copy only values explicitly present and return null for a missing value.” Check the schema and every extracted value. Record invented details and deviations from the requested format. -
Test the local setup. Run the coding prompt through each local frontend or runtime you can operate. Record the model revision, precision, backend, device, peak memory, elapsed time, output changes and failures. Repeat with network access disabled and record whether the model loads, completes and produces the same result.
Which rows of the comparison matter
In the Claude vs ChatGPT vs Gemini matrix, read the rows for Free plan, Low-cost tier, Main paid plan, Heavy-use plans, Team plan, Context window in the app, How usage limits are described, Training on your chats and Ads.
Those rows help separate hosted subscriptions from self-hosting, but they do not settle model licensing, downloadable weights, hardware requirements or installation work. Each matrix figure carries its vendor source and read date; an unconfirmed figure is marked “Not verified” and left blank, as explained in Chat Picker’s method. Statistics and test results from the older page that lacked sources have been removed.