
DeepSeek V4 Pro shows a market in which higher API prices can accompany a push for dependable, higher-value model work rather than a retreat from broad usage. Beginning in August 2026, DeepSeek raised API prices across products by about 3 to 12 times, according to woshipm. Its V4 Pro cache-hit input price moved from $0.003625 per million Tokens to $0.044, while output rose from $0.87 to $3.96.
The useful comparison is not a token rate alone, but the cost of finishing a task reliably.
Zhipu shows why higher prices need not suppress demand. geekpark reports that its open-platform and API business generated 825 million yuan in first-half revenue, with token usage growing more than 40-fold from the beginning of the year. The same account says average API selling prices rose about 101%. The pattern connects premium capability, cheaper broad-use access, and open weights as routes to a larger market.
Higher API prices can coexist with more usage
Zhipu's figures show why an API price increase does not automatically mean weaker demand. According to geekpark, its open-platform and API business produced 825 million yuan in first-half revenue, up 2735.7% year on year, and represented 86.5% of total revenue. Usage expanded at the same time.
- December 2025Annualized revenue was about $4 billion.
- March 2026Annualized revenue reached $6 billion.
- August 2026Annualized revenue had approached $13 billion.
By its interim results, Zhipu's MaaS platform had more than 40-fold token growth. That growth was from the beginning of the year.
Geekpark also reports that paid daily active users grew 603%. The average API selling price rose about 101%. Those movements complicate the familiar assumption that providers must continually cut token rates to win volume. A buyer may accept a higher listed rate when the model performs a valuable task reliably enough to justify repeated use. The relevant demand signal is therefore not price in isolation. It is paid activity alongside price.
DeepSeek provides a sharper example of the same tension, although woshipm's account concerns prices rather than usage. Beginning in August 2026, DeepSeek raised API prices. Different products rose by about 3 to 12 times. For DeepSeek V4 Pro, cache-hit input moved from $0.003625 per million Tokens to $0.044 per million Tokens. Output moved from $0.87 per million Tokens to $3.96 per million Tokens. Higher prices can coexist with a market in which capable API access remains worth buying for particular work.
Premium capability and cheap volume serve different jobs
Long-horizon coding and agent work create a different product tier from routine model use. Zhipu released GLM-5.2 in June. It focused on those demanding tasks, according to geekpark. The point is not merely that coding can command a premium; agent work asks a model to sustain progress across a longer process, where an early mistake can spoil the result. That makes reliability in completing the work more commercially relevant than a low token rate alone.
Lower-priced capacity can serve a separate job: broad, repeated use.
Zhipu's GLM-5.3-Flash illustrates that separation. geekpark reports that it uses a sparse architecture with 320 billion total parameters and 18 billion active parameters. It operates on a cluster of about 100,000 domestic chips. It costs about one-tenth as much as GLM-5.2. That is a route to cheaper volume without requiring the provider to position every model as the choice for long-horizon agents. Kimi K3's token cost is also about one-third of Anthropic Fable's, according to woshipm, reinforcing how sharply routine capacity can be priced below higher-cost alternatives.
Revenue figures show why the split matters. Zhipu's agent business generated 55.56 million yuan, up 304.4% year on year, geekpark reports. At MiniMax, open-platform and other AI enterprise-services revenue reached $73.93 million in the first half, up 703.1% year on year and representing 63.4% of total revenue; AI-native product revenue was $42.64 million, up 100.9% year on year. These figures point to demand for paid capability and enterprise access alongside low-cost model volume, rather than a market governed by one universal token price.

Open weights are an ecosystem acquisition channel
Kimi K3 and Qwen3.8-Max treat model weights as a route into other companies' products. According to woshipm, Kimi K3 can be downloaded, self-deployed, fine-tuned, and incorporated into a user's own product. Its license also allows modified versions to be published, distributed, sublicensed, or sold.
Alibaba likewise opened the underlying weights of Qwen3.8-Max, woshipm reports. That makes adoption easier for teams that want control over deployment or changes to the model rather than a standard hosted interface.
The free entry has a revenue boundary. A provider selling large-model-as-a-service built on Kimi K3 needs a separate Moonshot AI agreement once its revenue, including affiliates, exceeded $20 million during the prior 12 consecutive months. Qwen3.8 users face a separate commercial-license requirement above $50 million over the same period when they build model-as-a-service, AI coding, AI office assistants, or similar products. The weights can spread widely before the largest commercial uses become direct licensing opportunities.
That approach contrasts with a pure API-price strategy. geekpark says Zhipu's MaaS platform saw token usage grow more than 40-fold from the beginning of the year, while paid daily active users rose 603% and its average API selling price rose about 101%. Higher prices can therefore coexist with expanding demand when customers value the managed service.
Enterprise services offer another eventual revenue layer. According to geekpark, MiniMax generated $73.93 million from its open platform and other AI enterprise services in the first half, representing 63.4% of total revenue. It also said it had more than 2 million enterprise and developer customers. Open weights acquire users and builders; licenses and services seek payment later from the businesses that scale.
How Chinese model providers combine open access, workflow capability, and monetization
| Dimension | Kimi K3 | Qwen3.8-Max | Zhipu | MiniMax |
|---|---|---|---|---|
| Open-weight access | Full weights can be downloaded, self-deployed, fine-tuned, and used in users' own products. | Underlying weights are open; developers can download, deploy, and modify Qwen3.8 for free. | not covered | not covered |
| Commercial-use threshold | A separate agreement is required for large-model-as-a-service providers above $20 million in revenue over the previous 12 consecutive months. | A separate Qwen commercial license is required for specified commercial products above $50 million in revenue over the previous 12 consecutive months. | not covered | not covered |
| Workflow-oriented capability | Kimi's paper says overall capability remains below Claude Fable 5 and GPT-5.6 Sol, but is within the capability range of frontier models. | Alibaba Cloud's complete Qwen3.8-Max provides visual input, a default 1 million-Token context window, a thinking-mode toggle, and officially integrated tools. | GLM-5.2 focuses on long-horizon coding and agent capabilities. | not covered |
| Reported platform or enterprise revenue | not covered | not covered | Open-platform and API business generated 825 million yuan in first-half revenue and accounted for 86.5% of total revenue. | Open-platform and other AI enterprise-services revenue was $73.93 million in the first half and accounted for 63.4% of total revenue. |
| Reported pricing or unit-economics signal | Token cost is about one-third of Anthropic Fable's. | not covered | Average API selling price had risen about 101%; open-platform and API gross margin rose to 24.6%. | First-half gross margin rose from 12.1% to 17.9%. |
| Ecosystem or customer scale | not covered | In April 2025, global Qwen-series downloads exceeded 300 million and derivative models exceeded 100,000. | not covered | More than 2 million enterprise and developer customers. |
A download is not the same as a managed deployment
The downloadable Qwen3.8-2.4T-A95B weights are not a local copy of the complete Qwen3.8-Max service. According to woshipm, the download is a text-only model with a native 260,000-Token context window, extendable to about 1 million Tokens.
The files alone are nearly 4.9TB.
Alibaba Cloud's complete Qwen3.8-Max adds visual input, a default 1 million-Token context window, a thinking-mode toggle, and officially integrated tools, woshipm reports. Those features change what a buyer is deploying. Running downloaded weights can provide model access, but it does not automatically reproduce the hosted product's input handling or tool integration.
That distinction complicates every self-hosting break-even claim. The file size points to substantial deployment capacity before a request is served. A team must also operate the serving stack and decide who handles failures or performance changes. Comparing a cloud token bill with the apparent zero price of weights leaves those obligations outside the calculation.

The economics of managed capability are visible in provider financials, even when they do not yield a universal self-hosting price. geekpark reports that Zhipu recorded 252 million yuan in first-half gross profit at a 26.4% gross margin, while research and development spending reached 2.131 billion yuan, up 33.6% year on year. Hosted pricing can therefore reflect continuing work around a model, rather than inference alone.
Downloaded weights may reduce a license bill. They do not erase deployment work or support economics.
Choose by task completion path, not headline token price
- You need long-horizon coding or agent work, where completing a multistep workflow matters more than obtaining the lowest nominal token rate. Evaluate GLM-5.2 first: Zhipu released it in June with a focus on long-horizon coding and agent capabilities. Treat GLM-5.3-Flash as the lower-cost alternative only after testing whether its lower price meets the workflow's reliability requirements; it is priced at about one-tenth of GLM-5.2.
- You want to self-deploy, fine-tune, or embed a model in your own product rather than depend entirely on a hosted API. Consider Kimi K3 or Qwen3.8-Max weights. Kimi K3's weights can be downloaded, self-deployed, fine-tuned, and used in users' own products, while Qwen3.8 can be downloaded, deployed, and modified for free. Do not assume this removes commercial obligations: Kimi K3 requires a separate commercial agreement for certain large model-as-a-service providers whose revenue, together with affiliates, exceeded $20 million over the previous 12 consecutive months; Qwen requires a separate commercial license for specified product uses once revenue, together with affiliates, exceeded $50 million over the previous 12 consecutive months.
- Your application needs visual input, a default 1 million-Token context window, a thinking-mode toggle, or officially integrated tools. Use the complete Qwen3.8-Max on Alibaba Cloud rather than assuming the downloadable model is equivalent. The publicly downloadable Qwen3.8-2.4T-A95B weights are text-only, with a native context window of 260,000 Tokens that can be extended to about 1 million Tokens; the Alibaba Cloud offering adds the listed workflow features.
- You are selecting an open-weight model primarily to avoid infrastructure and operational costs. Do not equate downloadable with operationally lightweight. The Qwen3.8 model files on Hugging Face are nearly 4.9TB. Include storage, deployment, inference capacity, and licensing review in the cost of completing the task.
- You are buying API capacity for broad, high-volume usage and expect price stability. Model both usage growth and possible repricing rather than choosing solely on an introductory rate. Beginning in August 2026, DeepSeek raised API prices for different products by about 3 to 12 times; DeepSeek V4 Pro's cache-hit input price rose from $0.003625 per million Tokens to $0.044 per million Tokens, and its output price rose from $0.87 per million Tokens to $3.96 per million Tokens.
Buy a successful workflow rather than a token rate
Procurement should begin with a completed-task definition: which step is difficult, what output is acceptable, how long it may take, and who may handle the data. A token rate becomes useful only after those gates are set. The cheapest model that fails a critical handoff is not the lowest-cost workflow.
Route the hard step to the model that meets the acceptance test, then use cheaper or open options for routine transformations and overflow. Preserve a portable path where possible, so changing providers does not require redesigning the workflow. Measure task completion and latency at each handoff, rather than buying on benchmark reputation alone.
License scope is another gate, especially when an internal prototype becomes a product. According to woshipm, a company using Qwen3.8 for model-as-a-service, AI coding, AI office assistants, or similar products needs a separate Qwen commercial license when revenue with affiliates exceeded $50 million during the previous 12 consecutive months. The same outlet says Kimi K3 users in large-model-as-a-service need a separate Moonshot AI commercial agreement above $20 million on that basis.
The economics justify this discipline. geekpark reports that Zhipu's open-platform and API business produced 825 million yuan in first-half revenue and represented 86.5% of total revenue. Its gross margin moved from -0.4% to 24.6%, while revenue was up 2735.7% year on year. GLM-5.2, released in June, targets long-horizon coding and agent capabilities. Compare dependable completion with the full cost of the chosen route.
Treat the teams behind a model as an operational dependency. Track changes in research leadership, training capacity, and retention before committing a critical workflow to a provider. According to 36kr, Zhou Chang moved from Alibaba to ByteDance in the second half of 2024, while Cheng Ye'an joined Moonshot AI after contributing to GLM-4.5 and GLM-4.5V.
Talent movement can alter a roadmap faster than a token-price change.
Ask vendors who owns the relevant model work, how deployment support is staffed, and what happens if the team changes. ByteDance's reported loss of nearly 70 Seed-team technical employees is one signal to investigate; its later increases in bonuses and salary adjustments show why retention has become a competitive concern. For open-weight deployments, keep an exit path: document training, hosting, and evaluation choices so a change in a provider's people or priorities does not force a costly rebuild.
For readers outside China
- Availability: The source material establishes that Kimi K3 and Qwen3.8 weights are downloadable, and reports substantial overseas business for some Chinese providers: MiniMax generated $70.83 million in overseas revenue in the first half, representing 60.8% of total revenue. But it does not say which APIs, cloud products, payment methods, or support arrangements are available in any particular country outside China.
- Pricing: The sources provide only selective pricing. GLM-5.3-Flash is priced at about one-tenth of GLM-5.2. Kimi K3's Token cost is about one-third of Anthropic Fable's. Beginning in August 2026, DeepSeek raised API prices for different products by about 3 to 12 times; for DeepSeek V4 Pro, cache-hit input rose from $0.003625 per million Tokens to $0.044 per million Tokens and output rose from $0.87 per million Tokens to $3.96 per million Tokens. Current list prices for the other services are not disclosed in sources.
- Closest Western equivalents: Claude Fable 5, as a frontier-model comparison point for Kimi K3; GPT-5.6 Sol, as another frontier-model comparison point for Kimi K3; A commercial hosted API paired with downloadable open weights: the sources describe both routes for Kimi K3 and Qwen3.8, rather than presenting a single Western-style product equivalent
- Data residency: The source material does not cover data residency, processing location, retention, encryption, cross-border transfer, or enterprise compliance terms. Teams handling regulated or sensitive data should obtain those terms directly before sending production data to an API or cloud deployment.
Sources
- geekpark 智谱和 MiniMax,把大模型做成了两种生意 https://geekpark.net/news/369775
- woshipm 国内大模型开始走另一条路:模型免费下载,大客户另外谈钱 https://woshipm.com/ai/6459556.html
- woshipm 从40亿到130亿美元!中国大模型这波增长,可能和你想得不一样 https://woshipm.com/ai/6458981.html
- 36kr 从实习生日薪5000,到上亿年薪,AI人才争夺陷入疯狂丨深氪 https://36kr.com/p/3974476864876801
The evidence: 34 facts from 3 Chinese articles
Each line below was extracted from the article it sits under, in Chinese, before any of this was written. The writing is done from these and never from the source prose - that separation is structural, not a promise. How we work.
36kr从实习生日薪5000,到上亿年薪,AI人才争夺陷入疯狂丨深氪
- A colleague born in 1998 increased his annual salary from 1 million yuan to 3 million yuan after joining a leading major technology company.
- A postgraduate student born in the 2000s received a daily internship salary of 5,000 yuan before graduating.
- Zhou Chang, who had led multimodal pre-training at Alibaba, joined ByteDance in the second half of 2024.
- Former SenseTime research director Liu Yu had coordinated more than 4,000 GPUs for model training by 2024.
- ByteDance announced in an all-staff email at the end of 2025 that investment in bonuses and salary adjustments would increase by 35% and 1.5 times, respectively.
- ByteDance introduced a new job-level system at the end of 2025 that raised the salary ceiling for every level.
geekpark智谱和 MiniMax,把大模型做成了两种生意
- In the first half of 2026, MiniMax generated revenue of $117 million, up 283.1% year on year.
- In the first half of 2026, Zhipu generated revenue of 954 million yuan, up 399.7% year on year.
- Zhipu's first-half gross profit was 252 million yuan, with a gross margin of 26.4%.
- Zhipu's first-half research and development spending was 2.131 billion yuan, up 33.6% year on year.
- Zhipu's open-platform and API business generated 825 million yuan in first-half revenue, up 2735.7% year on year and accounting for 86.5% of total revenue.
- Zhipu's enterprise general-model-related revenue was 67.04 million yuan in the first half, down 54.6% year on year.
- Zhipu's open-platform and API business gross margin rose from -0.4% in the same period of the previous year to 24.6%.
- Zhipu reported a first-half loss of 2.072 billion yuan, an adjusted net loss of 1.964 billion yuan, and an operating loss of 2.147 billion yuan.
- Zhipu released GLM-5.2 in June, focusing on long-horizon coding and agent capabilities.
- Zhipu's agent business generated 55.56 million yuan in revenue, up 304.4% year on year.
- MiniMax's first-half gross profit was $20.81 million, up 464.8% year on year, and its gross margin rose from 12.1% to 17.9%.
- MiniMax's open-platform and other AI enterprise-services revenue was $73.93 million in the first half, up 703.1% year on year and accounting for 63.4% of total revenue.
- MiniMax's AI-native product revenue was $42.64 million in the first half, up 100.9% year on year.
- MiniMax generated $70.83 million in overseas revenue in the first half, representing 60.8% of total revenue, while mainland China generated $45.75 million, representing 39.2%.
- MiniMax's first-half research and development expenses were $297 million, up 138.8% year on year; its loss for the period was $358 million, while adjusted net loss was $293 million.
woshipm国内大模型开始走另一条路:模型免费下载,大客户另外谈钱
- Moonshot AI released Kimi K3 at the end of July.
- Kimi K3 has 2.8 trillion parameters, activates about 104 billion parameters per run, and has a 1 million Token context window.
- Kimi K3's full model weights can be downloaded, self-deployed, fine-tuned, and used in users' own products.
- The Kimi K3 license permits use, copying, modification, publication, distribution, sublicensing, and sale of derivative versions.
- A company that provides large-model-as-a-service using Kimi K3 and whose revenue, together with its affiliates, exceeded $20 million over the previous 12 consecutive months must first sign a separate commercial agreement with Moonshot AI.
- Alibaba released Qwen3.8-Max on August 3.
- Qwen3.8-Max has 2.4 trillion parameters, activates about 95 billion parameters per run, and supports a context window of up to about 1 million Tokens.
- Alibaba opened the underlying weights of its Max-level model Qwen3.8-Max.
- Developers can download, deploy, and modify Qwen3.8 for free.
- A company using Qwen3.8 for model-as-a-service, AI coding, AI office assistants, or similar products must obtain a separate Qwen commercial license if its revenue, together with its affiliates, exceeded $50 million over the previous 12 consecutive months.
- The publicly downloadable Qwen3.8-2.4T-A95B weights are for a text-only model with a native context window of 260,000 Tokens that can be extended to about 1 million Tokens.
- The complete Qwen3.8-Max on Alibaba Cloud provides visual input, a default 1 million-Token context window, a thinking-mode toggle, and officially integrated tools.
- The Qwen3.8 model files on Hugging Face are nearly 4.9TB.