Chat Picker

AI

On-Premise AI Deployment: Security and Performance

Official documentation identifies license, runtime, hardware, and data-path constraints for self-hosted AI, but it cannot predict whether your configuration will meet your own security and performance targets.

Sources checked 2 Oct 2026

For on-premise AI deployment, official pages settle a limited set of facts: model access and licenses, runtime defaults, supported hardware paths, quantization options, and whether prompts stay local or enter a cloud service. They do not show that one configuration will meet your security or performance targets, and Chat Picker has not tested ChatGPT, Claude, or Gemini for this use.

What the vendors document

These are the projects we read, not the whole field. Their documentation supports a configuration check, not a product ranking.

As read on October 1, 2026, Meta’s Llama model README says the model weights are licensed for researchers and commercial entities. Access requires license acceptance, and the download links are signed, expire after 24 hours or a set download limit, and can be requested again after certain errors. The page says full bf16 Llama 4 inference requires at least four GPUs. Its listed Scout configuration uses two 80 GB GPUs for FP8 or one 80 GB GPU for Int4. Meta also says FP8 or Int4 reduces memory use at the cost of minimal accuracy loss, a claim you should verify against your own workload.

The Ollama FAQ, checked on October 1, 2026, says Ollama does not see prompts or data during local operation. For cloud-hosted models, it processes prompts and responses to provide the service but says it does not store, log, or train on their content. It collects basic account information and limited usage metadata that excludes prompt and response content. The server binds to 127.0.0.1 on port 11434 by default, the context window defaults to 4,096 tokens, and each model handles one parallel request at a time by default. Disabling Ollama’s cloud features also removes access to cloud-hosted models and web search.

The llama.cpp README, read on October 1, 2026, describes a C/C++ implementation without dependencies. It documents source builds, prebuilt binaries, Docker, an OpenAI-compatible API server, and a built-in web interface. Its listed integer quantization levels are 1.5-bit, 2-bit, 3-bit, 4-bit, 5-bit, 6-bit, and 8-bit. It also documents CUDA, AMD support through HIP, MUSA, Vulkan, and SYCL backends, plus partial CPU-and-GPU acceleration when a model exceeds total VRAM. These are documented options, not performance measurements.

The hosted plan pages do not settle on-premise requirements. OpenAI’s pricing page, read on October 1, 2026, lists Free, Go, Plus, and Pro plans and says opt-out from training on chats is available on each. Anthropic’s pricing page says the same for Free, Pro, and Max, while Team is not trained on by default. Google’s US plans page lists consumer tiers and directs business accounts toward Gemini Enterprise, but we could not verify a chat-training setting from that page. The pages we read do not provide local installation instructions or hardware requirements for these hosted plans.

The accuracy guidance is equally important. OpenAI’s ChatGPT accuracy guidance says the assistant can produce incorrect or misleading output and may sound confident when wrong; it recommends treating responses as first drafts and checking important claims. Anthropic’s Claude guidance calls incorrect or misleading output hallucination, advises against relying on Claude as the only source of truth, and recommends checking original cited pages for missing context. Google’s Gemini sources guidance documents a Sources button for available related links, including public websites, uploaded files, and connected Workspace content. That page does not claim that the response is correct.

What the documentation cannot tell you

These pages do not provide a shared benchmark for your hardware, model, context length, and workload. They also do not establish your threat model or show what happens after cloud features, search, connected files, or external services are enabled.

Local deployment changes where computation can occur, but it is not by itself a complete security result. You still need to check network exposure, credentials, file permissions, logs, retained data, update procedures, and the behavior of every connected service.

Performance also depends on the complete configuration. Model size, numerical format, memory capacity, context length, concurrency, and workload all matter. Chat Picker has no quality, speed, accuracy, reliability, or benchmark results of its own and does not rank these options.

How to check it yourself

  1. Set a fixed workload. Create a packet of synthetic support tickets, each with a timestamp, service, severity, and short summary. Give the same packet to every configuration you test. Look for missing details, unsupported additions, and changes across repeated runs. Record the model file, runtime version, context length, numerical format, and every error.

  2. Map the data path. Use the packet with realistic but fake personal data. Inspect configuration files, outbound connections, logs, temporary files, model storage, search connections, and application permissions. Repeat the test with the network disabled. Record every destination, credential used, retained copy, and cleanup step you observe.

  3. Check supported answers. Attach the packet and use this prompt: Using only the attached support-ticket packet, list every record that mentions a failed payment. For each one, quote the timestamp and service exactly, then write "not found" if the packet does not support the answer. Look at whether quotations match the source and whether the assistant states uncertainty. Record each unsupported or incorrect claim by type.

  4. Compare hardware settings. Run the same packet with the baseline format and each reduced-precision format supported by your chosen project. Look at memory use, model loading time, time to first token, completion time, and changes in the answers. Record the exact CPU, GPU, RAM, VRAM, context, precision, runtime flags, and driver versions.

  5. Test concurrency and recovery. Use a small client against the documented API, submit a fixed burst of identical and mixed requests, and repeat it during a model load and after an idle period. Look for queueing, rejected requests, context errors, crashes, and recovery. Record the queue, parallel-request, context, memory, and logging settings.

Which rows of the comparison matter

On the Claude, ChatGPT, and Gemini comparison matrix, start with Context window, How usage limits are described, and Training on your chats. Context and usage descriptions expose workload assumptions, while the training row provides a data-governance signal.

If you are considering a hosted fallback, also check the Free plan, Main paid plan, Heavy-use plans, and Team plan rows. Do not treat those application limits as evidence of local memory capacity, throughput, or deployment security. Each matrix figure carries its vendor source and read date; an unconfirmed entry is marked “Not verified” and left blank, as described in Chat Picker’s method page.

Sources

For current prices and limits, see the dated comparison pages.

Compare assistants