Chat Picker

ChatGPT

Chinese AI Chat Models: What Their API Pages List

This page lists what the official API and pricing pages of Chinese-developed chat models publish about pricing, context length, and features, with no testing done by Chat Picker.

Sources checked 3 Oct 2026

The API and pricing pages for DeepSeek, Alibaba Cloud Model Studio, Moonshot Kimi, and Z.ai list the model versions, context lengths, per-token prices, and supported features such as vision, tool calls, and caching. Chat Picker has not tested these models for output quality, speed, or accuracy; this page records only what the vendors publish on their own pages, read on 2026-10-01. An earlier version of this page listed statistics and test results without sources, and those have been removed.

What the vendors document

DeepSeek's API docs name two models: DeepSeek-V4.1-Flash (deepseek-flash) and DeepSeek-V4-Pro-0813 (deepseek-v4-pro). Both list a context length of 1M tokens and a maximum output of 384K tokens, and both run in non-thinking or thinking mode, with thinking on by default. deepseek-flash supports vision; deepseek-v4-pro does not. Both support JSON output, tool calls, the Responses API, the Anthropic API, and Chat Prefix Completion in beta; beta FIM Completion works only in non-thinking mode.

Prices are quoted per 1 million tokens, as read on October 2026, and billing covers the total input and output tokens. For deepseek-flash at 1M tokens: cache-hit input costs $0.003 off-peak and $0.006 at peak, cache-miss input $0.15 off-peak and $0.3 at peak, and output $0.6 off-peak and $1.2 at peak. For deepseek-v4-pro: cache-hit input $0.022 off-peak and $0.044 at peak, cache-miss input $0.66 off-peak and $1.32 at peak, and output $1.98 off-peak and $3.96 at peak. Off-peak rates are half of peak. Peak hours are 01:00–04:00 and 06:00–10:00 UTC on weekdays that are not Chinese public holidays; all other hours are off-peak.

Alibaba Cloud Model Studio's models page describes capabilities, not prices. It offers Qwen and third-party models for text, image, audio, and video, and points to qwen3.8-omni-flash for audio or video analysis and text generation, qwen3.8-omni-flash-realtime for real-time audio or video conversations, and qwen3.5-omni-plus for non-real-time audio output. The Omni model handles text, images, audio, and video together; a decision model returns classification, yes-or-no answers, and scoring with probability distributions and confidence.

Moonshot's Kimi page lists four models. kimi-k3 charges $3.00 per 1M tokens for a 5-minute cache write, $6.00 for a 1-hour cache write, $0.30 for cached input, $3.00 for input, and $15.00 for output, with a context window of 1,048,576 tokens. kimi-k2.7-code charges $0.19 per 1M tokens for cache-hit input, $0.95 for cache-miss input, and $4.00 for output, context 262,144. kimi-k2.7-code-highspeed charges $0.38, $1.90, and $8.00 on the same basis. kimi-k2.6 charges $0.16, $0.95, and $4.00 with a 262,144-token context. Prices exclude taxes. The API auto-caches repeated request prefixes in 5-minute and 1-hour tiers (default 5 minutes); a cache hit refreshes the entry without an extra cache-write charge. File-content extraction and file-storage APIs are temporarily free when you only upload and extract a document. The page estimates one token as roughly three to four English characters.

Z.ai's pricing page states all prices in USD. GLM-5.3-Flash lists $0.15 input, $0.03 cached input, and $0.50 output per 1M tokens, with cached-input storage limited-time free. GLM-4.6V lists $0.3 input, $0.05 cached input, and $0.9 output. GLM-4.7-Flash and GLM-4.6V-Flash list input, cached input, cached-input storage, and output as Free. Separate tools carry their own prices: Web Search is $0.01 per use, GLM-Image is $0.015 per image, CogView-4 is $0.01, CogVideoX-3 is $0.2 per video, GLM-ASR-2512 is $0.03 per MTok (about $0.0024 per minute), the beta GLM Slide/Poster Agent is $0.7 per MTok, and General-Purpose Translation is $3 per MTok.

On data handling and accuracy, the vendor pages we read do not state training opt-out settings, data-retention terms, or hallucination disclaimers. The pages we read settle pricing, model names, context limits, and feature support; they do not settle output reliability.

What the documentation cannot tell you

The API pages tell you what each model can be billed for and which features exist, but they do not tell you how a model performs on your tasks. Output quality, factual accuracy, response speed, and behavior on edge cases are not published in the material we read. Chat Picker has not run quality, speed, accuracy, or benchmark tests on any of these assistants, so it has no scores or rankings to offer. Whether a model follows your instructions, handles your language mix, or stays within the stated context limit on a real workload is something only your own trial shows.

How to check it yourself

Use this plan to record what each vendor actually does for your use case. For every step, write down the model name, the input you sent, what the answer contained, and the billed tokens if shown.

  1. DeepSeek pricing in practice: send the same long prompt off-peak and at peak. Look at the cache-hit versus cache-miss cost on your bill and whether thinking mode changed the token count. Write down the per-call difference.
  2. Alibaba Qwen multimodal: give qwen3.8-omni-flash an image for a structured description and a video for extracted text. Look at whether the output matches the listed capability. Write down what worked and what did not.
  3. Kimi context and caching: upload a long document to kimi-k3 (context 1,048,576) and ask about its whole content. Look at whether a cache write was billed and whether repeats lowered input cost. Write down the cache tier used.
  4. Z.ai free and paid tiers: run one request on GLM-4.7-Flash (Free) and one on GLM-5.3-Flash, then enable Web Search. Look at whether the free row met your need and note the $0.01 charge. Write down the model and the final bill.
  5. Cross-model fit: run one real task on one model from each vendor with the same prompt. Look only at whether the documented feature you needed appeared. Write down which documented capability each model used.

Which rows of the comparison matter

When you weigh these API pages against the assistants Chat Picker compares, the rows that carry over are context window, published per-token or per-plan price, free-tier availability, and any stated data-handling or training terms. The full matrix is on the Claude vs ChatGPT vs Gemini page, where each figure carries the vendor source and the date it was read; when a figure could not be confirmed on the day it was read, the site marks it "Not verified." Chat Picker's method for sourcing every figure is described on the method page.

Quick questions

Do these pages list request limits?

The vendor pages we read state context lengths and token budgets but no per-minute or per-day message caps for the API. Rate limits, if any, are not in the material we read.

Are there free tiers?

Z.ai lists GLM-4.7-Flash and GLM-4.6V-Flash with input, cached input, storage, and output all Free, and Kimi's file extraction and storage APIs are temporarily free for upload-and-extract use. DeepSeek and Alibaba pages we read do not list a free API tier.

Does the documentation cover data training?

The four vendor pages we read do not state whether chats or prompts are used for training or how data is retained. Confirm directly with each vendor before sending sensitive input.

Sources: - Models & Pricing | DeepSeek API Docs - Supported Models and Capabilities Overview - Alibaba Cloud Model Studio - Model Inference Pricing Explanation - Kimi API Platform - Pricing - Overview - Z.AI Developer Document

For current prices and limits, see the dated comparison pages.

Compare assistants