EastofSilicon Tools and workflows from the Chinese internet
AI & Agents

DeepSeek V4 Pro turns model routing into the real eval

15 min read 3,450 words 36krgeekparkwoshipm
Developer reviewing code on a laptop
A developer works through code on a laptop.Photo: Jakub Zerdzicki / Pexels

DeepSeek V4 Pro sits in a new evaluation pattern: the model card still matters, but the harder question is whether the right model can finish a tool-using job at an acceptable cost and speed.

woshipm points to the old scoreboard and the new pressure at once. DeepSeek's model card gives V4 Pro Max 93.5 on LiveCodeBench, 80.6% on SWE Bench Verified, 67.9 on Terminal Bench 2.0, 90.1 on GPQA Diamond, and 83.5 on the million-context MRCR test. DeepSeek V4 Pro also scores 55.4 on SWE Bench Pro. Those numbers are useful, but they do not decide which model should handle browser automation, financial analysis, or code repository management inside a live workflow.

That is where the benchmark target is moving. 36kr reports that the China Academy of Information and Communications Technology's Institute of Artificial Intelligence released a Trusted AI MCP special test covering six task categories, including web search, location navigation, and 3D design. Its focus is multi-tool collaboration, complex task execution, and real-environment interaction. In that frame, the real product is the router.

The eval target is the completed task, not the model card

The useful comparison is shifting from a model card to a finished job. According to 36kr, the China Academy of Information and Communications Technology's Institute of Artificial Intelligence released a Trusted AI benchmark report built around MCP, and the design choice matters: it tests models inside work that uses tools, interacts with environments, and requires multi-step execution.

DeepSeek V4 rollout and multimodal follow-on
  1. April 24, 2026DeepSeek releases V4 preview with V4 Pro and V4 Flash
  2. July 31DeepSeek officially updates V4 Flash
  3. August 21DeepSeek launches V4-Flash-Vision-Exp and multimodal APIs

Its six task categories are concrete: location navigation, web search, browser automation, financial analysis, code repository management, and 3D design. That is closer to the way agent products fail. A model can answer a prompt well and still break when a browser state changes, a repository action needs checking, or a tool result has to be folded back into the next step.

That is the frame Chinese coverage is moving toward.

The model-card numbers still matter, but they are becoming inputs rather than verdicts. Woshipm cites DeepSeek's card for V4 Pro Max at 93.5 on LiveCodeBench, 80.6% on SWE Bench Verified, 67.9 on Terminal Bench 2.0, 90.1 on GPQA Diamond, and 83.5 on the million-context MRCR test. The same coverage records DeepSeek V4 Pro at 55.4 on SWE Bench Pro, and V4 Pro Max at 57.9 on SimpleQA Verified while Gemini 3.1 Pro scored 75.6 in DeepSeek's comparison table.

The caution is visible in the gaps between early claims and reproducible checks. Woshipm says V4 Pro preview scored 72.1 on Terminal Bench 2.1, the official V4 Pro scored 87.9, and Claude Fable 5 scored 88.0. It also records DeepSeek's DeepSWE rising from 12.8 to 62.7, while Claude Fable 5 scored 70. The woshipm author estimates an unreleased benchmark-chart score at about 56-58 and argues that V4 Pro is quasi-Fable-level in Agent and Coding. Useful, but not final.

Price, speed, and context pull different work to different models

API tables invite a simple comparison, but the woshipm author's own routing choices show why that shortcut breaks. DeepSeek V4 Pro is listed at $0.43 per million input tokens and $0.87 per million output tokens, while its cached-hit input price is $0.0036 per million tokens. Claude Fable 5 sits at $10 per million input tokens and $50 per million output tokens. Grok 4.6 starts at $2 per million input tokens and $6 per million output tokens. The cheapest line item is obvious. The cheapest completed task is not.

Latency changes the answer.

Grok 4.5 is priced at $2 per million input Tokens and $6 per million output Tokens, according to woshipm, and its output speed is about 80 Tokens per second. It also supports text and image input, has a 500,000 Token context, and has entered Grok Build, Cursor, and the SpaceXAI developer platform, with office-tool integration for Word, PowerPoint, and Excel. That makes it easier to justify when the task needs those surfaces, even if another model has a lower raw token price.

Context length pulls work the other way. DeepSeek V4 Pro supports a 1 million Token context, and DeepSeek's materials say that under a 1 million Token condition, V4 Pro's single-Token inference computation is about 27% of V3.2's, with KV Cache about 10% of V3.2's. Those figures matter for long documents and repeated context, where cache behavior can dominate the bill.

Reasoning settings add another routing lever. DeepSeek V4 Pro separates reasoning into Non Think, Think High, and Think Max, so the same model can be run differently depending on retry risk and verification need. The woshipm author frames the broader constraint as a difficult triangle of performance, price, and speed.

DeepSeek home page with the main product entry points
DeepSeek's home page shows its main chat and model entry points.Screenshot: deepseek.com

That is why their planned usage is split, not consolidated. They expect to keep ChatGPT and Codex for the most difficult ultra-large tasks, daily office tasks, creative work, content production, browser control, and computer control, while using DeepSeek V4 Pro for many automation tasks, asynchronous tasks, and some automated Agent work in AIHOT. The practical unit is the routed job, not the winning model name.

Grok 4.6 shows why routing needs evidence gates

Grok 4.6 is the useful warning label on the routing story. According to woshipm, several overseas financial media outlets reported an August 12 release and framed Grok 4.6 as a model for long-running agents and complex tasks. The same woshipm check found that SpaceXAI's news page, model documentation, and developer update records still listed Grok 4.5 as the latest official entry at the author's verification time.

That split matters more than the nameplate.

A router can treat the report as a lead, not as a result. If Grok 4.6 appears to finish development work quickly, it belongs on the candidate list for agent coding. It does not yet earn a precise win-loss claim against DeepSeek V4 Pro, Grok 4.5, or any other model until the version is documented and the task can be repeated by another operator under the same settings.

The woshipm piece gives reasons to care. It says DeepSeek V4 Pro and Grok 4.6 official versions were released within 2 hours of each other overnight, with DeepSeek V4 Pro at 1.6T parameters and Grok 4.6 at 1.5T parameters. It also reports Grok 4.6 API starting pricing at $2 per million input tokens and $6 per million output tokens. Those are routing inputs: model size, availability timing, and token cost.

But the experience claims need a gate. The woshipm author used Grok 4.6 for about 4 consecutive hours of development and almost did not encounter tasks longer than 20 minutes. The author also said some tasks ended in 3 minutes without quality reduction, then paid $300 for a SuperGrok Heavy membership and planned to use Grok Build and Grok 4.6 as main daily Vibe Coding products.

That is a practitioner signal, not a benchmark, and it should be treated as one workflow note rather than a frontier verdict. It says a workflow may be worth trying. It does not prove a general frontier shift by itself.

Grok 4.5 sharpens the point.

It was released in July and positioned by SpaceXAI for programming, agent tasks, and knowledge work. If the official surface still points there while field reports point to Grok 4.6, a serious router should separate four questions: which model was actually called, what task was run, how long it took, and what verification caught. Elon Musk's claim that Grok 4.7 will surpass all models belongs outside that gate until it becomes runnable evidence.

Vision and reasoning settings change the bill before the model changes

DeepSeek V4 Pro is a useful reminder that "best model" is not the same as "right endpoint." According to woshipm, V4 Pro does not have multimodal capability. For image work, DeepSeek's relevant switch is DeepSeek-V4-Flash-Vision-Exp, which geekpark says was launched on August 21 together with multimodal API services.

That changes the routing question before any quality comparison starts.

Developers call the vision model by setting `model='deepseek-v4-flash-vision-exp'`. Geekpark reports three image input methods: base64 inline input, external URL input, and Files API input. Files API launched alongside the model, the interface itself is free of charge, and it lets developers upload images, receive a `file_id`, then reference that `file_id` in later requests. This is plumbing, but it is also cost control.

The billing detail is sharper than the branding. DeepSeek says one image in DeepSeek-V4-Flash-Vision-Exp occupies at most 384 tokens and carries no extra visual-processing premium beyond token billing. Geekpark says GPT and Claude usually consume 800 to 1100 tokens on images of the same resolution. DeepSeek also says the vision model is priced the same as V4-Flash, matches the official V4-Flash version on pure-text benchmarks such as Agent, reasoning, and world knowledge, and has multimodal Agent capability close to Opus-4.8.

xAI home page with product and model information
xAI's home page shows the company's model and product landing area.Screenshot: x.ai

Geekpark's examples show why this belongs in operations, not marketing. Its reporter fed DeepSeek Vision a close-up Niu Lai screenshot and asked for a pure-code AI remake; the SVG plus CSS version took 35 seconds and consumed 384 image tokens. A playable complete HTML game from the same screenshot took 36 seconds. A responsive HTML page from a landing-page design screenshot arrived after 26 seconds.

Reasoning settings can move the bill even harder. In geekpark's default thinking-mode test, 2001 completion tokens were all reasoning tokens and the actual output was zero. After setting `reasoning_effort: "none"`, the same test fell from 23 seconds to 7.9 seconds and produced a complete output. Geekpark had previously found thinking mode consumed more than 82% of the output budget when the official V4 Pro version was released. The router therefore needs to choose vision support and reasoning mode as first-class settings, before it chooses the model name.

Route by task shape, not by a single model ranking

  • You need a tool-using agent for web search, browser automation, finance analysis, code repository work, navigation, or 3D design. Use an MCP-style evaluation as the selection lens rather than a chat benchmark. The CAICT MCP special test explicitly focuses on multi-tool collaboration, complex task execution, and real-environment interaction across location navigation, web search, browser automation, financial analysis, code repository management, and 3D design.
  • You need low-cost, long-context, non-visual automation or asynchronous agent work. Consider DeepSeek V4 Pro. It supports a 1 million Token context, is priced at $0.43 per million input tokens and $0.87 per million output tokens, and the woshipm author planned to use it for many automation tasks, asynchronous tasks, and some automated Agent work. Do not choose it when the task requires images, because DeepSeek V4 Pro does not have multimodal capability.
  • You need image-to-code, screenshot understanding, or multimodal API work with predictable token accounting. Use DeepSeek-V4-Flash-Vision-Exp rather than DeepSeek V4 Pro. It supports base64 inline input, external URL input, and Files API input; it can be called with model='deepseek-v4-flash-vision-exp'; and DeepSeek says one image occupies at most 384 tokens and has no additional visual-processing premium. In GeekPark tests, it produced a responsive HTML page from a landing-page screenshot in 26 seconds and a playable complete HTML game from a Niu Lai screenshot in 36 seconds.
  • You need fast completion and the model's default thinking mode is burning the output budget. Turn off reasoning when verification needs are low. In GeekPark's test with default thinking mode enabled, 2001 completion tokens were all reasoning tokens and the actual output was zero; after setting reasoning_effort: "none", the same test shortened from 23 seconds to 7.9 seconds and produced a complete output.
  • You need local or consumer-PC deployment rather than a cloud-only API workflow. Look at smaller local models such as StartLux-V1.0-27B-Preview. It has 27B parameters, can currently run on consumer-grade personal computers, and ranked second overall in the MCP special test, above DeepSeek-V4-Flash. It also ranked first in location navigation and tied for first or ranked first alone in browser automation and financial analysis compared with the 1.6T-parameter DeepSeek-V4-Pro.

Local 27B agents belong in the router, with limits

StartLux belongs in the router because it changes the local-model slot from fallback chat to agent work. According to 36kr, StartLux-V1.0-27B-Preview was developed by Shanghai StartLux Science and Technology Co., Ltd., and can currently run on consumer-grade personal computers. That matters for privacy-sensitive jobs, offline work, and MCP-style tool use where sending every step to a cloud model is the wrong default.

The score is useful, but only if it is read narrowly.

In 36kr's MCP special test, StartLux-V1.0-27B-Preview ranked second overall with 27B parameters and scored 39.25. It ranked above DeepSeek-V4-Flash in the overall table and scored 5.34 percentage points higher than Qwen-3.6-27B at the same parameter scale. The test covered location navigation, web search, browser automation, financial analysis, code repository management, and 3D design, with 7 evaluation items in total. Some indicators and data referred to MCP-Universe, so the result points to agent behavior rather than general intelligence.

The more interesting comparison is not small versus large. It is task fit. 36kr says StartLux ranked first in location navigation and tied for first or ranked first alone in browser automation and financial analysis against DeepSeek-V4-Pro, whose total scale is 1.6 trillion parameters. Woshipm reports that DeepSeek V4 Pro activates about 49 billion parameters per Token and supports a 1 million Token context. Those are routing facts, not vanity specs.

StartLux used Qwen3.6-27B as its base model, with targeted post-training and an AI training AI method called Auto Research. 36kr describes it as China's first local Agent model to use Auto Research for post-training. The caveat is simple: local agents earn placement for constrained MCP work, while cloud models still carry the broader context, infrastructure, and task coverage.

Routing tradeoffs across models covered in Chinese AI sources

DeepSeek V4 ProDeepSeek-V4-Flash-Vision-ExpStartLux-V1.0-27B-PreviewGrok 4.5 / Grok 4.6Claude Fable 5
Primary fit described in sourcesAutomation tasks, asynchronous tasks, and some automated Agent work; positioned by the woshipm author as quasi-Fable-level in Agent and Coding fieldsMultimodal API services; image-to-code and image-to-game workflows tested by GeekParkLocal Agent model focused on multi-tool collaboration, complex task execution, and real-environment interaction in the MCP special testGrok 4.5 positioned for programming, agent tasks, and knowledge work; woshipm author plans to use Grok Build and Grok 4.6 as main daily Vibe Coding productsUsed as the comparison point for difficult coding and agent benchmarks; pricing and benchmark scores covered, but workflow fit otherwise not covered
ModalityDoes not have multimodal capabilitySupports image input through base64 inline input, external URL input, and Files API inputnot coveredGrok 4.5 supports text and image input; Grok 4.6 modality not coverednot covered
Context or image-token handlingSupports a 1 million Token contextOne image occupies at most 384 tokens and is billed by token without an additional visual-processing premium, according to DeepSeeknot coveredGrok 4.5 has a 500,000 Token contextnot covered
Parameter scale1.6 trillion total parameters; activates about 49 billion parameters for each Tokennot covered27B parametersGrok 4.6 has 1.5T parameters; Grok 4.5 parameter count not coverednot covered
Pricing covered in sources$0.43 per million input tokens and $0.87 per million output tokens; cached-hit input is $0.0036 per million tokensPriced the same as V4-Flash; exact V4-Flash price not disclosed in sourcesnot disclosed in sourcesGrok 4.5 official API price is $2 per million input Tokens and $6 per million output Tokens; Grok 4.6 API starting pricing is $2 per million input tokens and $6 per million output tokens$10 per million input tokens and $50 per million output tokens
Speed or latency evidenceUnder a 1 million Token condition, single-Token inference computation is about 27% of V3.2's, and KV Cache is about 10% of V3.2'sGeekPark tests: SVG + CSS remake took 35 seconds; HTML game took 36 seconds; landing page HTML took 26 seconds; disabling thinking shortened one test from 23 seconds to 7.9 secondsCan currently run on consumer-grade personal computersGrok 4.5 output speed is about 80 Tokens per second; woshipm author used Grok 4.6 for about 4 consecutive hours and almost did not encounter any tasks longer than 20 minutes; some tasks ended in 3 minutes without quality reduction, according to the authornot covered
Verification or reasoning controlsDivides reasoning into three levels: Non Think, Think High, and Think MaxGeekPark found default thinking mode produced 2001 completion tokens that were all reasoning tokens and zero actual output; setting reasoning_effort: "none" produced a complete outputUsed an AI training AI method called Auto Research during model trainingnot coverednot covered
Benchmark or test evidenceV4 Pro Max scored 93.5 on LiveCodeBench, solved 80.6% on SWE Bench Verified, scored 67.9 on Terminal Bench 2.0, scored 90.1 on GPQA Diamond, and scored 83.5 on the million-context MRCR test; V4 Pro scored 55.4 on SWE Bench ProDeepSeek says it matches V4-Flash on pure-text benchmarks such as Agent, reasoning, and world knowledge, and says its multimodal Agent capability is already close to Opus-4.8Second place overall in the MCP special test with an overall score of 39.25; first in location navigation; tied for first or ranked first alone in browser automation and financial analysis compared with DeepSeek-V4-Pronot covered for Grok 4.6; Grok 4.5 benchmark scores not coveredTerminal Bench 2.1 score 88.0; DeepSWE score 70; Gemini comparison table shows Claude Fable 5 pricing but broader benchmark coverage not covered
Deployment or integration constraintsAPI usage covered; woshipm author added 1000 yuan of credit to the DeepSeek APIDevelopers call it by setting model='deepseek-v4-flash-vision-exp'; Files API lets developers upload images to obtain a file_id and reference it laterCan run on consumer-grade personal computersGrok 4.5 has entered Grok Build, Cursor, and the SpaceXAI developer platform, and can integrate with Word, PowerPoint, and Excel; woshipm author paid $300 for a SuperGrok Heavy membership after testing Grok 4.6not covered

Treat the next eval as a routing drill. Put DeepSeek V4 Pro on agent and coding jobs where woshipm sees it as quasi-Fable-level: Terminal Bench 2.1 moved from 72.1 in preview to 87.9 official, close to Claude Fable 5 at 88.0, while DeepSWE rose from 12.8 to 62.7 against Fable 5 at 70.

Do not send it vision work. Woshipm states that DeepSeek V4 Pro has no multimodal capability.

Watch the clock as closely as the score. The woshipm author used Grok 4.6 for about 4 consecutive hours and almost did not hit tasks longer than 20 minutes; some ended in 3 minutes without quality loss. Keep a fast-model lane for that loop, and treat Musk's claim about Grok 4.7 surpassing all models as a thing to verify, not a routing rule.

For readers outside China

  • Availability: The source material covers Chinese developer-facing coverage and does not give a complete availability map outside China. DeepSeek-V4-Flash-Vision-Exp is described as available through multimodal API services and callable by developers with model='deepseek-v4-flash-vision-exp'. DeepSeek launched Files API alongside it. Grok 4.5 is described as available through Grok Build, Cursor, and the SpaceXAI developer platform. StartLux-V1.0-27B-Preview is described as able to run on consumer-grade personal computers, but public download, licensing, and overseas availability are not disclosed in sources.
  • Pricing: DeepSeek V4 Pro is priced at $0.43 per million input tokens and $0.87 per million output tokens, with cached-hit input priced at $0.0036 per million tokens. DeepSeek-V4-Flash-Vision-Exp is priced the same as V4-Flash; the exact V4-Flash price is not disclosed in sources. DeepSeek says one image in DeepSeek-V4-Flash-Vision-Exp occupies at most 384 tokens and is billed by token without an additional visual-processing premium. Grok 4.5 pricing is $2 per million input Tokens and $6 per million output Tokens. Grok 4.6 API starting pricing is $2 per million input tokens and $6 per million output tokens. Claude Fable 5 API pricing is $10 per million input tokens and $50 per million output tokens. The woshipm author paid $300 for a SuperGrok Heavy membership and added 1000 yuan of credit to the DeepSeek API, but those are user-reported spending examples, not general price schedules.
  • Closest Western equivalents: Claude Fable 5, used in the sources as a pricing and coding-agent benchmark comparison; Grok 4.5 and Grok 4.6, positioned in the sources around programming, agent tasks, knowledge work, and long-running complex tasks; GPT and Claude, cited by GeekPark as image-processing comparators that usually consume 800 to 1100 tokens for images of the same resolution; Gemini 3.1 Pro, cited in DeepSeek's comparison table on SimpleQA Verified
  • Data residency: Data residency, retention, enterprise isolation, and whether uploaded Files API images remain inside China are not disclosed in sources. The sources say Files API lets developers upload images to the platform to obtain a file_id and reference it in later requests, but they do not describe storage location, deletion behavior, or compliance guarantees.

Sources

The evidence: 69 facts from 4 Chinese articles

Each line below was extracted from the article it sits under, in Chinese, before any of this was written. The writing is done from these and never from the source prose - that separation is structural, not a promise. How we work.

36kr最前线|本地大模型 StartLux-27B 通过 MCP-Universe 测评:综合第二,多项专项第一

  • The China Academy of Information and Communications Technology's Institute of Artificial Intelligence released a test report for the Trusted AI large model benchmark MCP special test.
  • StartLux-V1.0-27B-Preview was developed by Shanghai StartLux Science and Technology Co., Ltd.
  • StartLux-V1.0-27B-Preview achieved second place in overall performance in the MCP special test with 27B parameters.
  • StartLux-V1.0-27B-Preview ranked above DeepSeek-V4-Flash in the MCP special test overall ranking.
  • The MCP special test covers six task categories: location navigation, web search, browser automation, financial analysis, code repository management, and 3D design.
  • The MCP special test includes 7 evaluation items, consisting of six specialized tasks and one comprehensive evaluation.
  • The MCP special test focuses on large models' performance in multi-tool collaboration, complex task execution, and real-environment interaction.
  • Some indicators and data in the MCP special test refer to the open-source project MCP-Universe.
  • The models tested included DeepSeek-V4-Pro (1.6T), DeepSeek-V4-Flash-0731 (284B), Step-3.7-Flash (198B), StartLux-27B-260715 (27B), Qwen-3.6-27B (27B), and AgentCPM-Explore (4B).
  • StartLux-V1.0-27B-Preview received an overall score of 39.25 in the MCP special test.
  • StartLux-V1.0-27B-Preview scored 5.34 percentage points higher than Qwen-3.6-27B among models with the same parameter scale.
  • StartLux-V1.0-27B-Preview ranked first in the location navigation task.
  • StartLux-V1.0-27B-Preview tied for first or ranked first alone in tasks including browser automation and financial analysis compared with the 1.6T-parameter DeepSeek-V4-Pro.
  • StartLux used Qwen3.6-27B as the base model for post-training targeted enhancement.
  • StartLux used an AI training AI method called Auto Research during model training.
  • Google launched the Gemma 4 series in April, and its 31B Dense version ranked third on an open-source model leaderboard.
  • Meta released and open-sourced the 30B local model Muse Glimmer in early August.
  • NVIDIA launched the open-source model Nemotron 3.5 Lightning in August, a 30B-parameter MoE model that can run directly on local devices such as RTX PCs.
  • StartLux-V1.0-27B-Preview can currently run on consumer-grade personal computers.

geekparkDeepSeek 上线多模态,我用它做了《牛来》小游戏|AI 上新

  • DeepSeek officially announced on August 21 that V4-Flash-Vision-Exp was launched and that multimodal API services were opened.
  • DeepSeek-V4-Flash-Vision-Exp can be called by developers by setting model='deepseek-v4-flash-vision-exp'.
  • DeepSeek-V4-Flash-Vision-Exp supports three image input methods: base64 inline input, external URL input, and Files API input.
  • DeepSeek-V4-Flash-Vision-Exp is priced the same as V4-Flash.
  • DeepSeek launched Files API alongside DeepSeek-V4-Flash-Vision-Exp, and the interface itself is free of charge.
  • Files API lets developers upload images to the platform to obtain a file_id and then reference the file_id in later requests.
  • GeekPark tested DeepSeek Vision by feeding it a close-up screenshot from Niu Lai and asking it to use pure code to draw an AI remake.
  • In GeekPark's Niu Lai image-to-code test, DeepSeek identified the cow's orange coloring, frowning eyebrows, thick lips, expression, and the subtitle 'Mom' at the bottom of the image.
  • GeekPark's Niu Lai SVG + CSS image remake test took 35 seconds and consumed 384 tokens for the image.
  • GeekPark tested DeepSeek by showing it a screenshot of two cows facing each other from Niu Lai and asking it to write an interactive parkour game based on the cow images.
  • DeepSeek generated a playable complete HTML game from the Niu Lai screenshot in 36 seconds.
  • GeekPark tested DeepSeek Vision by feeding it a screenshot of a payment product landing page design and asking it to output runnable HTML.
  • DeepSeek output a complete responsive HTML page 26 seconds after receiving the landing page design screenshot.
  • In GeekPark's test with default thinking mode enabled, 2001 completion tokens were all reasoning tokens and the actual output was zero.
  • After GeekPark set reasoning_effort: "none", the same test shortened from 23 seconds to 7.9 seconds and produced a complete output.
  • GeekPark previously found that thinking mode consumed more than 82% of the output budget when the official V4 Pro version was released.

woshipmDeepSeek V4 Pro 遇上 Grok 4.6,新一轮模型竞争开始 比什么?

  • DeepSeek V4 series released a preview version on April 24, 2026, including V4 Pro and V4 Flash.
  • DeepSeek officially updated V4 Flash on July 31 and stated that the official version of V4 Pro would be released later.
  • DeepSeek's website had not listed August 11 as a new first-release date for V4 Pro as of the woshipm author's verification time.
  • SpaceXAI's news page, model documentation, and developer update records still listed Grok 4.5 as the latest official entry as of the woshipm author's verification time.
  • DeepSeek V4 Pro is a mixture-of-experts model with 1.6 trillion total parameters.
  • DeepSeek V4 Pro activates about 49 billion parameters for each Token it processes.
  • DeepSeek V4 Pro supports a 1 million Token context.
  • DeepSeek V4 Pro retains the mixture-of-experts framework and multi-Token prediction method used by DeepSeek V3.
  • DeepSeek V4 Pro adds hybrid attention, constrained hyper-connections, and a new optimizer.
  • DeepSeek's technical materials state that under a 1 million Token condition, V4 Pro's single-Token inference computation is about 27% of V3.2's, and its KV Cache is about 10% of V3.2's.
  • DeepSeek V4 Pro divides reasoning into three levels: Non Think, Think High, and Think Max.
  • DeepSeek's model card reports that V4 Pro Max scored 93.5 on LiveCodeBench, solved 80.6% of tasks on SWE Bench Verified, and scored 67.9 on Terminal Bench 2.0.
  • DeepSeek's model card reports that V4 Pro Max scored 90.1 on GPQA Diamond and 83.5 on the million-context MRCR test.
  • DeepSeek V4 Pro scored 55.4 on SWE Bench Pro.
  • DeepSeek V4 Pro Max scored 57.9 on SimpleQA Verified, while Gemini 3.1 Pro scored 75.6 in the same DeepSeek comparison table.
  • Grok 4.5 was released in July and is positioned by SpaceXAI as a model for programming, agent tasks, and knowledge work.
  • Grok 4.5 supports text and image input and has a 500,000 Token context.
  • SpaceXAI's official API price for Grok 4.5 is $2 per million input Tokens and $6 per million output Tokens, with an output speed of about 80 Tokens per second.
  • Grok 4.5 has entered Grok Build, Cursor, and the SpaceXAI developer platform, and it can integrate with office tools such as Word, PowerPoint, and Excel.

woshipm一夜之间,DeepSeek V4 Pro和Grok 4.6把全球大模型推入了新的斩杀线。

  • DeepSeek V4 Pro official version and Grok 4.6 official version were released within 2 hours of each other overnight.
  • DeepSeek V4 Pro has 1.6T parameters, and Grok 4.6 has 1.5T parameters.
  • The woshipm author paid $300 for a SuperGrok Heavy membership after testing Grok 4.6.
  • The woshipm author added 1000 yuan of credit to the DeepSeek API after testing DeepSeek V4 Pro.
  • The woshipm author plans to use Grok Build and Grok 4.6 as his main daily Vibe Coding products.
  • The woshipm author plans to continue using ChatGPT and Codex for the most difficult ultra-large tasks, daily office tasks, creative work, content production, browser control, and computer control.
  • The woshipm author plans to use DeepSeek V4 Pro for many automation tasks, asynchronous tasks, and some automated Agent work in products such as AIHOT.
  • DeepSeek V4 Pro is priced at $0.43 per million input tokens and $0.87 per million output tokens.
  • DeepSeek V4 Pro cached-hit input is priced at $0.0036 per million tokens.
  • Claude Fable 5 API pricing is $10 per million input tokens and $50 per million output tokens.
  • Grok 4.6 API starting pricing is $2 per million input tokens and $6 per million output tokens.
  • DeepSeek V4 Pro preview scored 72.1 on Terminal Bench 2.1, DeepSeek V4 Pro official scored 87.9, and Claude Fable 5 scored 88.0.
  • DeepSeek's DeepSWE score rose from 12.8 to 62.7, while Claude Fable 5 scored 70.
  • DeepSeek V4 Pro does not have multimodal capability.
  • The woshipm author used Grok 4.6 for about 4 consecutive hours of development and almost did not encounter any tasks longer than 20 minutes.