Open-Source
Open-Source AI Chat Models: Llama, Mistral and ChatGLM
This guide explains the access approval for Llama 3.1, the open-weight status of Mistral Large 3, and the research and commercial terms for ChatGLM3-6B.
Sources checked 2 Oct 2026
Open-source and open-weight do not come with one uniform setup: the Llama models README requires license acceptance, Mistral Docs describes Mistral Large 3 as open-weight, and the ChatGLM3 README calls ChatGLM3-6B open-source. These pages document access, context, and hardware notes, but not which model will fit your workload best. Chat Picker has not tested these models; this is the list of projects we read, not the whole field.
What the vendors document
As read on October 1, 2026, the Llama models README lists Llama 3.1 in 8B, 70B, and 405B sizes, each with a 128K context length and TikToken-based tokenization. The same Llama models README says the weights are licensed for researchers and commercial entities after approval, and that an approved request covers all Llama 3.1 models and earlier versions. Signed download links are emailed and expire after 24 hours or a set download limit. The page says full-bf16 Llama 4 inference needs at least four GPUs; its Llama 4 Scout example lists two 80GB GPUs for FP8 or one 80GB GPU for Int4.
The Mistral Docs, read on October 1, 2026, call Mistral Large 3 an open-weight, general-purpose multimodal model, but do not state its weight-license terms. As read on October 1, 2026, the Mistral pricing page lists Free with limited messages and $10 per month in API credits, Pro at $14.99 per month excluding taxes, and Team at $24.99 per user per month excluding taxes. The same page lists Team storage of up to 30GB per user. These are hosted-plan terms, not a substitute for the model license.
The ChatGLM3 README, read on October 1, 2026, calls ChatGLM3-6B open-source and lists variants with 8K, 32K, and 128K sequence lengths. It says the weights are open for academic research and free commercial use is available after registration through a questionnaire. Its sample default-FP16 setup requires about 13GB of VRAM, while four-bit quantization causes some performance loss. CPU inference is supported but slower and requires about 32GB of memory; only local model loading is supported on macOS. The README also lists GLM-4-0520 and other GLM-4 variants on its API platform, but gives no self-hosting hardware terms for those entries.
For a self-hosted route, the Ollama FAQ, read on October 1, 2026, says local operation does not see users’ prompts or data. For cloud-hosted models, the Ollama FAQ says Ollama processes prompts and responses but does not store, log, or train on their content. Its default context is 4096 tokens, and disabling cloud features removes access to cloud-hosted models and web search.
The project pages do not supply a shared benchmark or accuracy guarantee. For hosted assistants, OpenAI’s accuracy note says ChatGPT can sound confidently wrong and should be checked against reliable sources. Anthropic’s accuracy note says not to rely on Claude as the only source of truth. Google’s Gemini source note says a Sources panel may appear when sources are available. These cautions do not establish the quality of a local model run.
What the documentation cannot tell you
Documentation establishes options, not results. It cannot show whether a model will follow your instructions, ground an answer correctly in your files, call tools safely, or run at an acceptable speed on your hardware. A listed context length also does not prove that your runtime, quantization, memory, and concurrency settings can sustain it.
License summaries do not settle your legal obligations. Record the exact model version and license, then review the full terms for your intended use. The earlier page’s unsourced statistics and test results have been removed.
A benchmark score is useful only when it identifies the model, quantization, runtime, hardware, prompts, and scoring method. A generic score cannot establish which option fits your workload.
How to check it yourself
Use the same documents and settings for every system, then record what happened rather than assuming an outcome.
-
Set a baseline. Give every system: “Create a launch checklist for a fictional support chatbot on a company laptop, covering setup, access control, monitoring, and rollback.” Look for explicit assumptions and unsupported claims; record the exact model, quantization, runtime settings, and date.
-
Test document grounding. Attach an access policy, support procedure, and sample ticket. Ask: “Answer the ticket using only the attached documents. Quote the supporting passage for each decision and say when the documents do not settle it.” Look for traceable support, omissions, and added facts; record unsupported statements, citation errors, and corrections.
-
Run three work tasks. Give each system: “Write a reply to the attached billing complaint using only its policy”; “Summarize the attached project memo and mark every uncertainty”; and “Repair this function and add tests:
def clean(rows): return {r['id']: r['name'].strip() for r in rows}.” Look for policy grounding, preserved uncertainty, and runnable code; record errors, omissions, and manual edits. -
Check tools and API output. Create
sample-config.jsoncontaining{"timeout":0,"retries":-1}. Ask: “Read it, identify invalid values, and return only JSON with keysvalid,errors, andproposed_changes.” Use a disposable folder; look for valid JSON and unrequested changes, and record setup failures or manual steps. -
Measure the runtime. Repeat one prompt with hardware, power settings, context, and parallelism fixed. Record model-file size, peak RAM or VRAM, load time, time to first token, response time, failures, and restarts. Keep these as setup observations, not a general ranking.
-
Review access and the data route. Describe the intended users, deployment location, model, and commercial purpose. Ask: “List questions to verify about model licensing, data handling, and operations; do not give a legal conclusion.” Check every item against the exact project documents; record approval, registration, local or cloud routing, retention terms, and unresolved gaps.
Which rows of the comparison matter
In the Claude vs ChatGPT vs Gemini matrix, focus on Free plan, Main paid plan, Context window in the app, How usage limits are described, and Training on your chats.
Those rows help compare hosted plans, but they do not answer weight access, licensing, runtime, hardware, quantization, or deployment data-path questions. Add separate notes for the exact open model, license or approval path, runtime, hardware, context settings, and whether inference stays local. Chat Picker’s method records a source and read date for each matrix figure and marks unconfirmed cells “Not verified”; it publishes no assistant quality, speed, or accuracy scores of its own.