
GLM 5.3 is a Zhipu coding model that makes the AI developer stack look less like a subscription choice and more like a replaceable toolchain. One account says Zhipu released it on August 14, then described GLM 5.3 as tied with Kimi K3 for first place among open-source coding models on benchmark rankings after release. The same account put it at the level of closed-source flagships including Claude Fable 5 and GPT-5.6 Sol.
That matters because the switching layer already exists.
woshipm cites CNBC reporting that many U.S. companies are leaving OpenAI and Anthropic for more cost-effective Chinese models such as DeepSeek, Zhipu, and Qwen. It also points to OpenRouter, a platform for accessing multiple AI models, where Chinese models have made up more than 30% of weekly token consumption since February. The useful question is no longer which model deserves loyalty. It is how coding teams set up editors, routers, agents, and review steps so the preferred model can change without breaking the work.
What GLM-5.3 changes in a coding setup
GLM-5.3 matters less as a declaration of allegiance than as a new slot in the coding stack. Zhipu released it on August 14, according to a Chinese tech outlet, after GLM-5.2 had already been described by ifanr as a top Chinese programming model. That sequence is the point: the model arrives as an upgrade candidate for a role developers already understand, not as a reason to rebuild every tool choice around one vendor.
- May 2025Zhipu holds emergency strategy meeting on next-stage models
- July 2025GLM 4.5 goes online after delayed launch and direction change
- September 2025Zhipu launches GLM Coding Plan
- February 2026Zhipu releases major model update GLM 5
- June 13, 2026Zhipu releases flagship model GLM 5.2
- August 14Zhipu releases GLM 5.3
Treat it as swappable by default.
The benchmark story supports that framing. The outlet said GLM-5.3 matched Kimi K3 for first place among open-source coding models after release, and also placed it at the same level as closed-source flagships including Claude Fable 5 and GPT-5.6 Sol. ifanr described comparable programming capability against Fable 5 and GPT-5.6 Sol, and noted that its chart put GLM-5.3 beside Kimi K3, Fable 5, and GPT-5.6 Sol. Those comparisons do not prove permanent superiority. They show that the coding-model seat is now contested by several models close enough to be evaluated task by task.
That is why the setup lesson is practical. A team can keep Claude, OpenAI, DeepSeek, Kimi, or Qwen in the conversation while still adding GLM-5.3 as a candidate for coding prompts. The question becomes: which model handles this repository, this refactor, this agent run? Not which logo should own the whole workflow.
OpenRouter makes that behavior visible at market scale. woshipm notes that OpenRouter lets developers access multiple AI models, and says Chinese models have accounted for more than 30% of weekly token consumption there since February. Their share once reached 46%. CNBC, as relayed by woshipm, also reported U.S. companies moving from OpenAI and Anthropic toward more cost-effective Chinese models such as DeepSeek, Zhipu, and Qwen. That does not make GLM-5.3 the final answer. It makes it part of a rotation where the winning coding model can change without breaking the workflow.
How to plug it into editors, routers, and agents
For an English-speaking developer, the most practical way to read GLM-5.3 is as another model endpoint in a stack that already expects rotation. woshipm describes OpenRouter as a platform for accessing multiple AI models, which matters because routers reduce model choice to configuration instead of a rewrite. If GLM-5.3 performs well on a coding task, it can sit beside other models rather than replace the whole setup.
Zhipu also has its own path. According to a Chinese tech outlet, it launched the GLM Coding Plan in September 2025 and became the first Chinese large model company to launch a Coding Plan. ifanr says GLM-5.3 is fully available to GLM Coding Plan users and is open for subscription.
The native tools are ZCode and AutoClaw. ifanr says GLM-5.3 is available in Zhipu's official agent application ZCode and productivity tool AutoClaw, while woshipm says both launched with GLM-5.3 on the release day. That gives developers a packaged route if they want the model, the coding surface, and the agent layer from one vendor.
The broader signal is integration breadth. woshipm lists early access in TraeWork, TraeCode, Coze, WorkBuddy, CodeBuddy, Qoder, QwenWork, CatPaw, JoyCode, and OpenCode. Those names matter less as endorsements than as connection points: GLM-5.3 is being routed into editors, agent builders, and coding assistants where teams already compare models task by task.
The maturity questions sit around the edges. ifanr says API access is expected to open next Tuesday, with full weights open-sourced within two weeks. woshipm also says GLM Coding Plan quota for all users was reset at 13:00 on the release day. For developers outside China, the deciding details will be documentation, billing, support, and English-language product flow.
Why price pressure favors model rotation
The useful reading of the CNBC-cited OpenRouter numbers is rotation, not replacement. According to woshipm, Chinese models made up only 4.5% of tokens consumed on OpenRouter in the first half of 2025 and averaged 11% over the past 12 months. Yet since February, their weekly share has been above 30%, and at one point it reached 46%.
That pattern looks like teams testing price ceilings rather than pledging allegiance.
The reason is blunt: inference bills are now large enough to change model choice. Woshipm says Chinese AI models are priced at 10% to 40% of comparable Anthropic and OpenAI models. On output-token pricing, GLM-5.2 costs about 18% as much as Opus 4.8. DeepSeek in thinking mode is priced at $0.435/$0.87, which woshipm compares with GPT-5.5 at $5/$30; its non-thinking mode is $0.14/$0.28. GLM-5.2 sits higher at $1.4/$4.4, but that is still about 28% of GPT-5.5's input price and about 14.7% of its output price.
This is why a single migration story is too neat. Woshipm cites CNBC saying many U.S. companies are leaving OpenAI and Anthropic for DeepSeek, Zhipu, and Qwen. It also cites a U.S. AI company founder who said on X that in early June he moved 100% of company traffic from Anthropic to DeepSeek V4, saving millions of dollars and improving performance in many use cases.
But model churn cuts both ways. Ifanr says internal test users on social media felt GLM-5.3 was better to use than Kimi K3 and DeepSeek V4-Pro-0813. Woshipm's author frames the market shift as a contest over whose unit intelligence is cheapest, not simply whose model is strongest. For developers, that means the durable skill is designing a coding stack where the preferred model can change without forcing a rewrite of the workflow.
Why post-training matters more than brand loyalty
Zhipu's model strategy moved in two layers. According to a Chinese tech outlet, the company held an emergency strategy meeting in May 2025 and decided to stop treating text, multimodality, and coding as separate vertical tracks. The new direction was a large-parameter model fused around Reasoning, Coding, and Agentic data. That shift delayed GLM 4.5 from its original April schedule until July 2025, backed by 15 trillion tokens of general data and 8 trillion tokens of Coding, Reasoning, and Agentic data.

GLM-5.3 makes the quieter point: the base does not have to change for the product to move.
ifanr reported that GLM-5.3 stayed at roughly the same 700 billion parameter level and used a 743B base model. woshipm also noted that its base model was unchanged from GLM-5.2. Zhipu's own technical account, as summarized by ifanr, attributed the gains mainly to post-training scaling, not a fresh foundation-model leap. That is why Slime matters. Zhipu open-sourced the post-training framework, says it has used it since GLM 4.5, and says it covers the workflow needed for release-level post-training. It currently supports the GLM series, the Qwen series, some DeepSeek models, and Llama 3.
The benchmark changes show why brand loyalty is a weak operating rule. woshipm recorded a 50% improvement over GLM-5.2 on Zhipu's internal coding evaluation. Terminal-Bench 3.0 rose from 4.6 to 28.3, DeepSWE v1.1 from 46.2 to 66.9, and Agents' Last Exam from 23.8 to 28.5. GLM-5.3 also scored 1769 on GDPval-AA v2, which covers 44 occupations.
But teams should not outsource judgment to benchmark names. On Z.ai Code Bench in High mode, GLM-5.3 reached 31.4% accuracy, while Claude Opus 4.8 reached 29.5% in its highest mode. GLM-5.3 used about 50,000 tokens per task on average; Claude Opus 4.8 used about 120,000. The useful question is how those tradeoffs behave inside your repositories.
Coding-model stack signals across GLM-5.3, DeepSeek, OpenRouter, and Claude/OpenAI references
| Dimension | GLM-5.3 / Zhipu | DeepSeek / other Chinese models | Claude / OpenAI references | Workflow takeaway |
|---|---|---|---|---|
| Coding benchmark position | Matched Kimi K3 for first place among open-source models in coding on benchmark rankings after release; benchmark results showed programming capability comparable to Fable 5 and GPT-5.6 Sol | DeepSeek-V4-ProMax was very close to top U.S. models on LiveCodeBench, Terminal Bench, and BrowseComp, and exceeded Claude Opus4.6 on some metrics | Claude Fable 5, GPT-5.6 Sol, Claude Opus4.6, GPT-5.4, and Claude Opus 4.8 are used as reference points across sources | Do not hard-code loyalty to one frontier model; keep benchmarks and routing swappable |
| Agentic workflow capability | AutomationBench tests cross-application workflow orchestration capability; Agents' Last Exam tests agent capability; GLM-5.3's Agents' Last Exam score increased from 23.8 to 28.5 | not covered | not covered | Agent use cases should be evaluated separately from ordinary coding prompts |
| Engineering-agent efficiency | On Z.ai Code Bench in High mode, GLM-5.3 reached 31.4% accuracy, while Claude Opus 4.8 reached 29.5% in its highest mode; GLM-5.3 output about 50,000 tokens per task on average, while Claude Opus 4.8 needed about 120,000 tokens | not covered | Claude Opus 4.8 is the comparison model in the cited Z.ai Code Bench rows | Model choice can affect both answer quality and token budget per task |
| Security-review capability | GLM-5.3 scored 84.5% on CyberGym; 54.4% on ExploitBench; on ExploitGym completed 105 vulnerability exploitation tasks within two hours and 130 tasks within six hours | not covered | Mythos 5 scored 83.8% on CyberGym, 78.0% on ExploitBench, and completed 181 vulnerability exploitation tasks within two hours and 247 tasks within six hours; GPT-5.6 Sol scored 83.6% on CyberGym and 76.5% on ExploitBench | First-pass security review is becoming part of the coding-model stack, not a separate afterthought |
| Pricing and cost pressure | GLM-5.2 is priced at $1.4/$4.4, about 28% of GPT-5.5's input price and about 14.7% of GPT-5.5's output price; GLM-5.3 pricing not disclosed in sources | Chinese AI models are priced at 10% to 40% of comparable Anthropic and OpenAI models; DeepSeek is priced at $0.435/$0.87 in thinking mode and $0.14/$0.28 in non-thinking mode | OpenAI GPT-5.5 is priced at $5/$30, and Anthropic Fable5 is priced at $10/$50 | Cost-tested routing matters because the cheapest adequate model may change |
| Adoption and switching evidence | GLM-5.3 is available in Zhipu's official agent application Zcode and productivity tool AutoClaw; GLM-5.3 is fully available to GLM Coding Plan users and is open for subscription | A U.S. AI company founder said in early June he had switched 100% of his company's traffic from Anthropic to DeepSeek V4, saving millions of dollars and improving performance in many use cases | CNBC reported that many U.S. companies are abandoning OpenAI and Anthropic and switching to more cost-effective Chinese models | Production setups need abstraction layers so traffic can move when price-performance changes |
| Multi-model access layer | TraeWork, TraeCode, Coze, WorkBuddy, CodeBuddy, Qoder, QwenWork, CatPaw, JoyCode, and OpenCode opened early access for GLM-5.3 | OpenRouter lets developers access multiple AI models; since February, Chinese AI models have accounted for more than 30% of weekly token consumption on OpenRouter and once rose as high as 46% | OpenRouter is framed as an alternative access layer to single-vendor use | Use platforms and tooling that make model replacement operationally easy |
| Open-source and post-training infrastructure | Zhipu trained GLM-5.3 using a 743B base model; GLM-5.3 has roughly the same parameter scale as the previous version, at the 700 billion parameter level; Zhipu open-sourced Slime; Slime supports GLM, Qwen, some DeepSeek models, and Llama 3 | Some DeepSeek models are supported by Slime | not covered | Post-training and tooling compatibility may matter as much as the base model |
How to place GLM-5.3 in a swappable AI coding stack
- You are optimizing for coding cost rather than loyalty to a single frontier provider. Treat GLM-5.3 and peer Chinese models as candidates in a routing layer, not as a one-time replacement. Sources describe developers using OpenRouter to access multiple AI models, and Chinese models have accounted for more than 30% of weekly token consumption on OpenRouter since February, at one point reaching 46%. Pricing evidence also supports cost-testing: GLM-5.2 is priced at $1.4/$4.4, while OpenAI GPT-5.5 is $5/$30 and Anthropic Fable5 is $10/$50.
- You need a coding model for agentic software-engineering tasks, terminal work, or long-cycle engineering agents. Include GLM-5.3 in evaluation. It improved Terminal-Bench 3.0 from 4.6 to 28.3, DeepSWE v1.1 from 46.2 to 66.9, and Agents' Last Exam from 23.8 to 28.5. On Z.ai Code Bench in High mode, GLM-5.3 reached 31.4% accuracy, while Claude Opus 4.8 reached 29.5% in its highest mode, and GLM-5.3 used about 50,000 tokens per task on average compared with about 120,000 tokens for Claude Opus 4.8.
- You want first-pass security review or vulnerability-discovery assistance, but not to hand off final security judgment to a model. Use GLM-5.3 as a screening and triage component, while retaining human security review. Sources report that GLM-5.3 scored 84.5% on CyberGym and 54.4% on ExploitBench, and completed 105 vulnerability exploitation tasks within two hours and 130 tasks within six hours on ExploitGym. But the same sources show Mythos 5 completed 181 tasks within two hours and 247 tasks within six hours on ExploitGym, so GLM-5.3 is not uniformly ahead of specialist frontier competitors.
- You need stable production API access immediately after release. Do not assume day-one API availability. One source says the GLM 5.3 API was originally scheduled to go online on August 18 but was delayed. Another says API access was expected to open next Tuesday, and the full weights would be open-sourced within two weeks. Build fallback routes to other models rather than hard-wiring GLM-5.3 into a production path before access is confirmed.
- You are choosing tools for hands-on use rather than calling the model directly. Start with the integrations named in the sources: GLM-5.3 is available in Zhipu's official agent application Zcode and productivity tool AutoClaw, and ZCode and AutoClaw launched with GLM-5.3 on the release day. TraeWork, TraeCode, Coze, WorkBuddy, CodeBuddy, Qoder, QwenWork, CatPaw, JoyCode, and OpenCode opened early access for GLM-5.3.
Where security review fits in the workflow
Security is where the model-rotation story needs the most restraint. ifanr says GLM-5.3 performed on par with Claude Mythos 5 in cybersecurity-related tasks, but the detailed rows split that claim into different jobs. On CyberGym, woshipm reports GLM-5.3 at 84.5%, ahead of GLM-5.2 at 77.2%, Mythos 5 at 83.8%, and GPT-5.6 Sol at 83.6%. That supports a parity argument for some defensive or assessment-style work.
Exploit generation is a different test.
On ExploitBench, GLM-5.3 scored 54.4%, while Mythos 5 reached 78.0% and GPT-5.6 Sol reached 76.5%, according to woshipm. ExploitGym shows the same gap in a more operational form: GLM-5.3 completed 105 vulnerability exploitation tasks within two hours and 130 within six hours, while Mythos 5 completed 181 within two hours and 247 within six hours. GLM-5.2 was far behind both, at 29 within two hours and 39 within six hours.
That makes GLM-5.3 easier to place in a workflow. It can help with first-pass review: scan a diff, flag suspicious patterns, summarize risk, and prepare issues for a human security owner. It should not be treated as a drop-in exploit engineer just because a broad cybersecurity score looks competitive.
Zhipu's own rollout points in that direction. woshipm says the company began investing in cybersecurity research in September 2025 during GLM-5.2 and GLM-5.3 development, and plans to open GLM-5.3 weights two weeks after release, after security evaluation and model hardening. ifanr also notes a public disclosure ledger for vulnerabilities found by its models.
The permission model should match that maturity curve. Let an agent read proprietary code under tight boundaries before it gets write access. Require review before Bash commands that change files or touch networks. Keep auto-approval for narrow, reversible checks, not for exploit chains or production repositories.
The promise is real, but bounded. Since GLM-5.2, Zhipu and security teams found 2,436 vulnerabilities after screening and deduplication, including 1,097 medium- and high-risk vulnerabilities; woshipm says Zhipu estimated their value at 30 million yuan using public prices from Zerodium, Crowdfense, Apple Security Bounty, and Pwn2Own.
Treat GLM-5.3 as a candidate in your coding stack, not as a new default. Put it behind the same router or agent harness you use for Claude, OpenAI, and other models, then compare it on your own repositories. The open thread is security-sensitive automation: woshipm says Zhipu began investing in cybersecurity research in September 2025 during GLM-5.2 and GLM-5.3 development, and the jump shows up in task results.
Keep permissions tight.
GLM-5.3 scored 84.5% on CyberGym and 54.4% on ExploitBench, while GLM-5.2 scored 77.2% and 24.4%. On ExploitGym, GLM-5.3 completed 105 exploitation tasks within two hours and 130 within six hours; GLM-5.2 completed 29 and 39. That makes it useful for first-pass code audit, but also a reason to watch Bash access, auto-approval, and tool execution logs before letting any agent act freely.
Make model-assisted review a normal pre-merge step. Woshipm says Zhipu and security teams found 2,436 vulnerabilities after initial screening and deduplication, including 1,097 medium- and high-risk vulnerabilities, and estimated their value at 30 million yuan using public bounty prices.
For readers outside China
- Availability: GLM-5.3 is described as available in Zhipu's official agent application Zcode and productivity tool AutoClaw, and as fully available to GLM Coding Plan users and open for subscription. The sources also say API access was expected to open next Tuesday and that full weights would be open-sourced within two weeks; another source says the API had been scheduled for August 18 but was delayed. Availability outside China is not disclosed in sources.
- Pricing: The sources provide pricing for GLM-5.2, not GLM-5.3. GLM-5.2 is priced at $1.4/$4.4, about 28% of GPT-5.5's input price and about 14.7% of GPT-5.5's output price. OpenAI GPT-5.5 is priced at $5/$30, Anthropic Fable5 at $10/$50, DeepSeek at $0.435/$0.87 in thinking mode, and DeepSeek at $0.14/$0.28 in non-thinking mode. GLM-5.3 subscription or API pricing is not disclosed in sources.
- Closest Western equivalents: Claude Opus 4.8 for long-cycle coding-agent comparison; Anthropic Fable5 for flagship-model pricing and capability comparison; OpenAI GPT-5.5 for pricing comparison; GPT-5.6 Sol for benchmark comparison; Claude Mythos 5 for cybersecurity-task comparison
- Data residency: The source material does not cover data residency, regional hosting, enterprise data controls, or whether prompts and code sent through Zhipu tools are processed inside or outside China. It also does not disclose compliance posture for non-Chinese enterprise customers.
Sources
- 36kr 一个扭转命运的决定:5000 亿智谱如何「逆风改命」|深氪 https://36kr.com/p/3947451908357257
- ifanr 实测GLM-5.3: 在神仙打架的一周杀回国模顶流,还按下了重置键 https://ifanr.com/1675225
- woshipm 智谱和DeepSeek,用性价比攻破美国大门? https://woshipm.com/ai/6427542.html
- woshipm 智谱GLM-5.3发布:前沿编程能力与涌现的网络安全能力 https://woshipm.com/ai/6447066.html
The evidence: 60 facts from 4 Chinese articles
Each line below was extracted from the article it sits under, in Chinese, before any of this was written. The writing is done from these and never from the source prose - that separation is structural, not a promise. How we work.
36kr一个扭转命运的决定:5000 亿智谱如何「逆风改命」|深氪
- Zhipu released GLM 5.3 on August 14.
- The GLM 5.3 API was originally scheduled to go online on August 18 but was delayed.
- Zhipu released the major model update GLM 5 in February 2026.
- Zhipu was founded in 2019.
- Zhipu's market capitalization once reached 1.3 trillion Hong Kong dollars.
- Zhipu held an emergency strategy meeting in May 2025 to discuss the direction of its next-stage models.
- Zhipu decided at the May 2025 strategy meeting to shift from separate vertical models for text, multimodality, and coding to a large-parameter three-in-one model combining Reasoning, Coding, and Agentic data.
- Zhipu's 2025 financial report showed a net loss of 4.718 billion yuan and research and development spending of 3.182 billion yuan.
- GLM 4.5 went online in July 2025 after its launch was delayed from the original April schedule because Zhipu temporarily adjusted the model direction.
- Zhipu prepared 15 trillion tokens of general data and 8 trillion tokens of Coding, Reasoning, and Agentic data to train GLM 4.5.
- Zhipu launched the GLM Coding Plan in September 2025 and became the first Chinese large model company to launch a Coding Plan.
- Zhipu listed in Hong Kong in January and CEO Zhang Peng gave each employee a 200 yuan red envelope on the listing day.
- Zhipu released the flagship model GLM 5.2 on June 13, 2026.
ifanr实测GLM-5.3: 在神仙打架的一周杀回国模顶流,还按下了重置键
- Zhipu released GLM-5.3 after GLM-5.2 had been described as a top Chinese programming model.
- GLM-5.3 was compared with Kimi K3, Fable 5, and GPT-5.6 Sol in the visualization chart.
- AutomationBench tests cross-application workflow orchestration capability.
- Agents' Last Exam tests agent capability.
- GLM-5.3 has roughly the same parameter scale as the previous version, at the 700 billion parameter level.
- Zhipu trained GLM-5.3 using a 743B base model.
- Zhipu open-sourced a post-training framework called Slime.
- Slime currently supports post-training for the GLM series, the Qwen series, some models in the DeepSeek series, and Llama 3.
- In the ExploitGym cybersecurity test, GLM-5.3 completed 130 of 898 questions within a 6-hour time limit.
- Zhipu publicly released a cybersecurity disclosure ledger that records vulnerabilities found by its models in different projects.
- GLM-5.3 is available in Zhipu's official agent application Zcode and productivity tool AutoClaw.
- GLM-5.3 is fully available to GLM Coding Plan users and is open for subscription.
- GLM-5.3 API access is expected to open next Tuesday, and the model's full weights will be open-sourced within two weeks.
woshipm智谱和DeepSeek,用性价比攻破美国大门?
- CNBC reported that many U.S. companies are abandoning OpenAI and Anthropic and switching to more cost-effective Chinese models such as DeepSeek, Zhipu, and Qwen.
- OpenRouter is a platform that lets developers access multiple AI models.
- OpenRouter data cited by CNBC showed that Chinese models accounted for only 4.5% of tokens consumed on the platform in the first half of 2025.
- OpenRouter data cited by CNBC showed that Chinese models averaged only 11% of tokens consumed on the platform over the past 12 months.
- Since February, Chinese AI models have accounted for more than 30% of weekly token consumption on OpenRouter.
- Chinese AI models' share of token consumption on OpenRouter once rose as high as 46%.
- The founder of a U.S. AI company said on X that in early June he had switched 100% of his company's traffic from Anthropic to DeepSeek V4.
- DeepSeek-V4-ProMax was compared with mainstream Chinese and foreign models at release on LiveCodeBench, Terminal Bench, and BrowseComp.
- Zhipu's GLM-5.1Thinking scored 58.4 on SWE Pro, higher than the roughly 57 level of Claude Opus 4.6 and GPT-5.4.
- Zhipu's official data showed that GLM-5.2 differed from Claude Opus4.8 by about 1% on the FrontierSWE long-cycle engineering agent benchmark.
- Based on output token pricing, GLM-5.2 costs about 18% as much as Opus 4.8.
- OpenAI GPT-5.5 is priced at $5/$30, and Anthropic Fable5 is priced at $10/$50.
- DeepSeek is priced at $0.435/$0.87 in thinking mode, about 8.7% of GPT-5.5's input price and about 2.9% of GPT-5.5's output price.
- DeepSeek is priced at $0.14/$0.28 in non-thinking mode.
- GLM-5.2 is priced at $1.4/$4.4, about 28% of GPT-5.5's input price and about 14.7% of GPT-5.5's output price.
woshipm智谱GLM-5.3发布:前沿编程能力与涌现的网络安全能力
- Zhipu released GLM-5.3, whose base model is unchanged from GLM-5.2.
- GLM-5.3 improved by 50% over GLM-5.2 in Zhipu's internally built coding evaluation.
- Zhipu will open GLM-5.3 model weights two weeks after release, after completing security evaluation and model hardening.
- Zhipu's official coding tool ZCode and productivity tool AutoClaw launched with GLM-5.3 on the release day.
- TraeWork, TraeCode, Coze, WorkBuddy, CodeBuddy, Qoder, QwenWork, CatPaw, JoyCode, and OpenCode opened early access for GLM-5.3.
- GLM Coding Plan quota for all users was reset at 13:00 on the release day.
- GLM-5.3's Terminal-Bench 3.0 score increased from 4.6 to 28.3.
- GLM-5.3's DeepSWE v1.1 score increased from 46.2 to 66.9.
- GLM-5.3's Agents' Last Exam score increased from 23.8 to 28.5.
- GLM-5.3 scored 1769 on GDPval-AA v2, which covers 44 occupations.
- On Z.ai Code Bench in High mode, GLM-5.3 reached 31.4% accuracy, while Claude Opus 4.8 reached 29.5% in its highest mode.
- On Z.ai Code Bench, GLM-5.3 output about 50,000 tokens per task on average, while Claude Opus 4.8 needed about 120,000 tokens.
- Zhipu began investing in cybersecurity research in September 2025 during the development of GLM-5.2 and GLM-5.3.
- GLM-5.3 scored 84.5% on CyberGym, compared with GLM-5.2's 77.2%, Mythos 5's 83.8%, and GPT-5.6 Sol's 83.6%.
- GLM-5.3 scored 54.4% on ExploitBench, compared with GLM-5.2's 24.4%, Mythos 5's 78.0%, and GPT-5.6 Sol's 76.5%.
- On ExploitGym, GLM-5.3 completed 105 vulnerability exploitation tasks within two hours and 130 tasks within six hours.
- On ExploitGym, GLM-5.2 completed 29 vulnerability exploitation tasks within two hours and 39 tasks within six hours.
- On ExploitGym, Mythos 5 completed 181 vulnerability exploitation tasks within two hours and 247 tasks within six hours.
- Since the release of GLM-5.2, Zhipu and security teams found 2,436 vulnerabilities after initial screening and deduplication, including 1,097 medium- and high-risk vulnerabilities.