EastofSilicon Tools and workflows from the Chinese internet
AI & Agents

GLM-5.3 turns coding-model choice into a systems test

11 min read 2,466 words 36krifanrwoshipm
Rows of server racks with accelerator hardware
Server racks house the accelerator hardware used to run large AI models.Photo: panumas nikhomkhai / Pexels

GLM-5.3 makes coding-model choice a systems test: it scored 60 alongside Kimi K3 despite having 744B parameters against Kimi K3's 2.8T. woshipm reports that GLM-5.3 ranked first among Chinese models in the latest Artificial Analysis ranking, while Kimi K3 and GLM-5.3 placed fourth globally.

The comparison is not a simple parameter contest. Zhipu's official blog, as cited by woshipm, says GLM-5.3 shares GLM-5.2's base model and attributes its gains to post-training; it also claims about 50% improvement on its proprietary Code Bench. ifanr's Tang Jie sees something narrower but consequential: a model can optimize a system that, in turn, serves the model. That loop shifts attention from impressive single outputs toward completed work, the tools and infrastructure behind it, and the cost of reaching a reliable result.

Post-training sets a starting point, not a purchase decision

GLM-5.3's 40B active parameters sit inside a 744B-parameter model, a distinction that already complicates simple model shopping. woshipm reports that GLM-5.3 and Kimi K3 each scored 60 in Artificial Analysis, despite Kimi K3 carrying 2.8T parameters. The same ranking placed the pair fourth globally and first among Chinese models.

Sunrise's spin-off, chip release, and financing sequence
  1. the end of 2024SenseTime spins off its chip business as Sunrise
  2. January 2026Sunrise releases its S3 inference chip
  3. April 2026Sunrise completes a financing round led by Hangzhou Capital
  4. August 2026Sunrise's new financing round reportedly includes multiple investors

Those figures make parameter count a weak shortcut.

According to woshipm, Zhipu said GLM-5.3 uses the same base model as GLM-5.2, with all gains coming from post-training. Zhipu also reported an improvement of about 50% on its proprietary Code Bench and described GLM-5.3 as the strongest open-weight coding model available. That is meaningful evidence of a better starting point, especially because it separates post-training effects from a new base model. It does not settle which model completes coding work best under a fixed operating setup.

woshipm describes a controlled comparison of Kimi K3, GLM-5.3, and DeepSeek-v4-pro-0813. Each model received the same task types, prompts, scoring standards, official APIs, and Claude Code. The author's judgment across the four tasks put K3 first overall, GLM-5.3 next, and DeepSeek-v4-pro far behind. That ordering differs from a shared score in Artificial Analysis and from a parameter-count comparison. Benchmarks remain useful for identifying capability, but controlled task execution reveals how a model behaves when the coding environment is held constant.

Tang Jie, writing for ifanr, argues that GLM-5.3 is still far from recursive self-improvement. Yet he sees a minimal loop emerging: a model improves a system, and that system serves the model. Post-training begins that story; it cannot finish the purchase decision.

A finished agent task exposes the gaps hidden by demos

A polished first attempt can conceal the difference between a convincing demo and a task that survives its own interaction. The evaluation described by woshipm covered frontend webpages, 3D scenes, websites, and games. Its scene prompt asked for a Three.js webpage showing 100,000 heavenly soldiers attacking Huaguo Mountain, with the plot, viewpoints, and settings aligned to *Journey to the West*.

K3 and GLM-5.3 each completed that scene in one attempt. The woshipm author judged GLM-5.3's result ahead of K3's, with DeepSeek-v4-pro-0813 behind both. That is meaningful visual evidence, but it does not establish that the model can carry a user through a longer chain of stateful work.

The game task makes the gap visible. Models had to create a 3D interactive web game from the script file "源记-游戏.md". Only K3's version was playable from beginning to completion. GLM-5.3's controls failed in the second level, while DeepSeek-v4-pro errored in the first level and then jumped unexpectedly to the second level.

Completion changes the ranking. For the game, woshipm judged K3 ahead of GLM-5.3, again ahead of DeepSeek-v4-pro-0813. The most complex game also took more than 10 minutes for DeepSeek-v4-pro, 43 minutes for GLM-5.3, and 67 minutes for K3. Practitioners therefore need measurements absent from a one-attempt result: how an agent recovers after a broken control, and whether the same task completes reproducibly. Latency must be read alongside those outcomes, especially where long-context and multimodal deployment already faces memory capacity and interconnect-bandwidth limits.

GLM-5.3's infrastructure agent closes a narrow feedback loop

GLM-5.3-Flash offers a compact example of an infrastructure agent working inside a serving stack rather than merely generating code snippets. According to ifanr, the model moved from its first run on domestic AI accelerators to carrying all production traffic in two weeks. End-to-end throughput rose by 3.2 times during that deployment process.

The claimed loop is narrow, but concrete.

ifanr reports that a GLM-5.3-powered Infra Agent helped examine bottlenecks, suggest optimization plans, and make some code changes. In KDA's context-parallel path, it identified accumulated TF32 rounding errors during state-matrix merging as the source of a long-context precision problem. The fix entered Flash Linear Attention project PR #1180.

Another finding concerned KV Transfer and DeepEP Dispatch. Their ineffective overlap meant transfer overhead exceeded 30% in some cases; after a change that promptly released the Python GIL on the intra-node path, that additional overhead fell below 1%. The agent also improved a Decode Kernel by 1.71 times after removing four repeated executions of the same normalization calculations.

This is more persuasive than a generic claim that an agent can optimize inference. Each result links diagnosis to a code-level intervention and a measured serving outcome. Yet the account does not establish that such changes can be safely granted production access, or that outside teams can independently reproduce the gains.

The Z.ai home page for its AI services and coding tools
Z.ai's home page presents its AI services and coding tools.Screenshot: z.ai

woshipm supplies adjacent evidence of technical analysis: its author reports that GLM-5.3 found six vulnerabilities in an open-source Markdown editor, and records an 84.5% CyberGym score against 83.8% for Mythos 5 and 83.6% for GPT-5.6 Sol. Those results do not demonstrate recursive improvement. As Tang Jie argues, the meaningful development is a minimal loop: the model helps optimize a system that then serves the model.

The three layers shaping practical AI coding performance

GLM-5.3 post-training and coding modelGLM-5.3-Flash inference systemSunrise S3 inference chip
Primary roleText-only model for coding and other tasks; it does not have visual recognition capability.Production inference system for GLM-5.3-Flash.Inference-specific chip architecture.
Reported improvement pathZhipu's official blog stated that GLM-5.3 uses the same base model as GLM-5.2 and that every gain comes from post-training.Optimization included ReplaySSM, intra-node tensor parallelism, INT8, FP8 and BF16 mixed-precision caching, and Encode-Prefill-Decode separation.Unlike mainstream GPUs, S3 does not pursue a unified training-and-inference design and concentrates transistor and power budgets on inference efficiency.
Evidence on coding-task outputThe evaluation judged it behind K3 but ahead of DeepSeek-v4-pro-0813 overall across four coding tasks; it produced a 3D scene in one attempt.Not covered.Not covered.
Long-context or memory constraintNot covered.Deployment had to support 1 million-token long context while addressing domestic chip memory-capacity and interconnect-bandwidth limits.Uses LPDDR memory rather than the HBM typically used by unified training-and-inference GPUs.
Measured or reported performanceScored 84.5% on CyberGym; GLM-5.3 and Kimi K3 both scored 60 on Artificial Analysis.End-to-end throughput increased by 3.2 times during deployment; end-to-end service performance improved by about 3 times compared with its initial version.The S3 technical combination still needs to prove itself through large-scale delivery and a declining cost curve per million tokens, according to the woshipm author.
Self-improving operational loopA GLM-5.3-powered Infra Agent analyzed bottlenecks, proposed optimization plans, and completed some code changes.The Infra Agent improved a Decode Kernel by 1.71 times and reduced additional KV Transfer overhead from more than 30% to below 1% after a modification.Not covered.
Cost informationAPI pricing: 8 yuan per million input tokens, 2 yuan per million cached input tokens, and 28 yuan per million output tokens.Zhipu claims hardware utilization and per-token cost were close to mainstream NVIDIA GPU platforms.Not disclosed in sources.
Scale or deployment statusAvailable through Z.ai, the Zhipu Qingyan app, the BigModel Open Platform, and GLM Coding Plan.Deployed on a cluster of more than 100,000 domestic AI accelerators; it carried all production traffic in two weeks.Released in January 2026.

Inference cost depends on memory and software together

Sunrise's S3 takes a hardware-first route to inference cost. According to a Chinese technology outlet, it abandons the unified training-and-inference design common to mainstream GPUs, concentrating transistor and power budgets on inference efficiency. Its use of LPDDR rather than HBM makes memory architecture part of the buying decision, especially when long contexts keep accumulated state resident.

Zhipu's GLM-5.3-Flash instead shows how software can stretch domestic accelerator deployments. ifanr reports that the system ran across more than 100,000 domestic AI accelerators while addressing memory-capacity limits, interconnect bandwidth, multimodal requests, and 1 million-token context support. Its stack combines ReplaySSM, intra-node tensor parallelism, INT8, FP8 and BF16 mixed-precision caching, plus Encode-Prefill-Decode separation. Zhipu claims utilization and per-token cost close to mainstream NVIDIA GPU platforms.

The memory choice does not remove the software problem.

Agent work makes KV-transfer behavior especially consequential because one deliverable can trigger repeated long-context calls. ifanr found scenarios where KV Transfer failed to overlap effectively with DeepEP Dispatch, pushing transfer overhead above 30%. Promptly releasing the Python GIL on the intra-node path cut additional KV Transfer overhead to below 1%. The Infra Agent also raised a Decode Kernel's performance by 1.71 times after removing four repeated normalization calculations. That evidence shifts evaluation from chip labels toward the cost and latency of completed work. Sunrise's inference-specific design may lower the memory-side burden, but its technical combination still has to demonstrate large-scale delivery and a declining cost curve per million tokens.

Choose for the task loop, not the parameter count

  • You need a coding model for long-context or high-volume agent workflows, where repeated calls can make serving performance a practical constraint. Treat inference infrastructure as part of the model choice. GLM-5.3-Flash was adapted for 1 million-token long-context support and multimodal requests, and its end-to-end throughput increased by 3.2 times during its two-week deployment process. Zhipu also reports end-to-end service performance improved by about 3 times from the initial version.
  • You are comparing models mainly by total parameter count. Use task results and post-training evidence alongside size. GLM-5.3 and Kimi K3 both scored 60 on Artificial Analysis, despite GLM-5.3 having 744B parameters and Kimi K3 having 2.8T parameters. Zhipu said GLM-5.3 uses the same base model as GLM-5.2 and that its gains come from post-training.
  • You need one-shot generation for a self-contained frontend or 3D web scene. GLM-5.3 is a plausible candidate to test against Kimi K3. In an evaluation using the same task types, prompts, scoring standards, official APIs, and Claude Code, K3 and GLM-5.3 each produced the 3D scene in one attempt; the evaluator judged GLM-5.3 better than K3 for that task.
  • You need an agent to finish a complex, multi-level interactive game rather than merely generate an initial prototype. Do not assume strong benchmark or single-shot performance guarantees completion. In the game-development task, only K3's version was playable from beginning to completion; GLM-5.3 had nonfunctional controls in the second level. The reported run time was 43 minutes for GLM-5.3 and 67 minutes for K3, while DeepSeek-v4-pro-0813 took more than 10 minutes.
  • You are operating a large inference fleet on domestic accelerators and need to reduce bottlenecks rather than only change models. Prioritize an instrumentation-and-optimization loop. GLM-5.3-Flash was deployed on more than 100,000 domestic AI accelerators, with constraints including memory capacity and interconnect bandwidth. Its Infra Agent identified a KDA long-context precision problem and a KV Transfer issue; after an intra-node change, additional KV Transfer overhead fell from more than 30% to below 1%.

Build a scorecard around completed work

A coding-model scorecard should begin with completed work in one agent environment, not a comparison of model labels. Run GLM-5.3, Kimi K3, and DeepSeek-v4-pro-0813 against identical task types, prompts, scoring standards, official APIs, and Claude Code-the setup used in the woshipm evaluation.

Measure task completion first, then inspect output quality. Record end-to-end latency, retries, failed tool calls, context consumption, and the price of each successful task. A model that produces an impressive draft but cannot finish the task has not earned its apparent speed.

The differences can be concrete. woshipm found that only K3 completed a playable game from beginning to completion; GLM-5.3's second-level controls did not work, while DeepSeek-v4-pro-0813 failed in the first level and jumped unexpectedly to the second. In the most complex game case, DeepSeek-v4-pro-0813 took more than 10 minutes, GLM-5.3 took 43 minutes, and K3 took 67 minutes. The woshipm author ranked K3 first overall, GLM-5.3 second, and DeepSeek-v4-pro-0813 well behind.

Price needs the same task-level treatment. GLM-5.3 carries the same API pricing as GLM-5.2: 8 yuan per million input tokens, 2 yuan per million cached input tokens, and 28 yuan per million output tokens, according to woshipm. Those rates matter only alongside the context and retry record.

Treat infrastructure actions as a separate risk gate. ifanr reports that GLM-5.3's Infra Agent can analyze bottlenecks, propose optimization plans, and complete some code changes. Require human review, sandbox execution, rollback paths, and access controls before such changes reach production. Zhipu's claim that GLM-5.3-Flash approaches mainstream NVIDIA GPU platforms on utilization and per-token cost is worth testing against that controlled scorecard.

Run your own completion test before assigning a model recurring agent work. A model that produces a convincing game build is not necessarily one that survives the full path: according to woshipm, GLM-5.3's controls failed in the second level, while K3 alone completed the game from start to finish.

Time the whole job, including repair attempts and tool failures. In woshipm's complex game case, DeepSeek-v4-pro took more than 10 minutes, GLM-5.3 took 43 minutes, and K3 took 67 minutes. That gap makes elapsed time a poor proxy for usable output. Test the routes, state changes, and recovery points your users actually need. Also separate visual tasks from text tasks: GLM-5.3 is text-only and has no visual recognition capability. For each candidate, keep the task artifact and record where execution breaks.

For readers outside China

  • Availability: GLM-5.3 is available through Z.ai, the Zhipu Qingyan app, the BigModel Open Platform, and GLM Coding Plan. Its availability outside China is not disclosed in sources. GLM-5.3 ran anonymously as Ox-Alpha on OpenCode and OpenRouter, but the source material does not state whether it remains available there under its disclosed name.
  • Pricing: GLM-5.3 has the same API pricing as GLM-5.2: 8 yuan per million input tokens, 2 yuan per million cached input tokens, and 28 yuan per million output tokens. Pricing for Kimi K3 and DeepSeek-v4-pro-0813 is not disclosed in sources.
  • Closest Western equivalents: Claude Code, which was used as the coding harness in the cited comparison; GPT-5.6 Sol, which appears as a CyberGym comparison point; Claude Opus 5, which is mentioned as a rumored 10T-scale model
  • Data residency: The source material does not cover data residency, regional processing locations, retention, enterprise privacy controls, or cross-border data-transfer terms for Z.ai, the Zhipu Qingyan app, the BigModel Open Platform, GLM Coding Plan, OpenCode, or OpenRouter.

Sources

The evidence: 54 facts from 4 Chinese articles

Each line below was extracted from the article it sits under, in Chinese, before any of this was written. The writing is done from these and never from the source prose - that separation is structural, not a promise. How we work.

36kr曦望再融20亿元,四个月估值翻倍至200亿 | 独家

  • Sunrise recently raised a new round of 2 billion yuan, reaching a post-investment valuation of about 20 billion yuan.
  • In April 2026, Sunrise completed a financing round of more than 1 billion yuan led by Hangzhou Capital, with a valuation exceeding 10 billion yuan.
  • Sunrise's valuation nearly doubled in less than six months between its April 2026 financing and its recent financing.
  • Since being spun off from SenseTime at the end of 2024, Sunrise has cumulatively raised nearly 6 billion yuan.
  • Sunrise declined to comment further on its reported financing.
  • At the end of 2024, SenseTime spun off its high-investment, long-cycle chip business as part of its "1+X" organizational restructuring.
  • SenseTime co-founder Xu Bing became chairman of Sunrise after the chip business was spun off.
  • Sunrise co-CEO Wang Yong previously worked as a core architect at AMD and Kunlunxin and joined SenseTime in 2020 to lead its chip business.
  • Sunrise co-CEO Wang Zhan previously served as a Baidu vice president.
  • Before its spin-off, the Sunrise team mass-produced two chips: the multimodal visual inference chip S1 and the large-model inference GPU S2.
  • In Sunrise's new financing round in August 2026, investors reportedly included Charoen Pokphand Group, Andon Health, Infore Environment Technology Group, Tongcheng Travel, PICC Equity Investment, CCB Equity Investment, Janchor Partners, CAS Star and Cowin Capital.
  • Sunrise released its S3 inference chip in January 2026.
  • Unlike mainstream GPUs, S3 does not pursue a unified training-and-inference design and instead concentrates its transistor and power budgets on inference efficiency.
  • S3 uses LPDDR memory rather than the HBM typically used by unified training-and-inference GPUs.
  • Some LPDDR5X products from ChangXin Memory Technologies entered mass production in 2025, and its next-generation LPDDR6 development is nearing completion.
  • Moore Threads, MetaX and Biren Technology have listed since the end of 2025, while Enflame Technology and Kunlunxin are among companies preparing to list.

ifanr刚刚,智谱首个 RSI 成果发布,10 万国产卡用 GLM 造 GLM

  • Zhipu GLM chief scientist Tang Jie shared research on GLM-5.3-Flash inference-system optimization on X.
  • GLM-5.3-Flash went from its first run on domestic AI accelerators to carrying all production traffic in two weeks.
  • The system's end-to-end throughput increased by 3.2 times during the two-week deployment process.
  • A GLM-5.3-powered Infra Agent participated in optimization work for the GLM-5.3-Flash inference system.
  • GLM-5.3-Flash was deployed on a cluster of more than 100,000 domestic AI accelerators.
  • The deployment had to address domestic chip memory-capacity and interconnect-bandwidth limits, new-model-architecture adaptation, 1 million-token long-context support, and multimodal request processing.
  • The GLM-5.3-powered Infra Agent helped analyze performance bottlenecks, propose optimization plans, and complete some code changes.
  • Zhipu's optimization measures included ReplaySSM, intra-node tensor parallelism, INT8, FP8 and BF16 mixed-precision caching, and an Encode-Prefill-Decode separation architecture.
  • GLM-5.3-Flash's end-to-end service performance improved by about 3 times compared with its initial version.
  • GLM-5.3-Flash ran under the anonymous model name Ox-Alpha on OpenCode and OpenRouter.
  • Ox-Alpha became one of the most-used models on both OpenCode and OpenRouter within one week and processed more than 62 trillion tokens in six days.
  • The Infra Agent identified a long-context calculation-precision issue in KDA's context-parallel path caused by accumulated TF32 rounding errors during state-matrix merging.
  • The fix for the KDA precision issue was merged into Flash Linear Attention project PR #1180.
  • The Infra Agent found that KV Transfer did not overlap effectively with DeepEP Dispatch, causing transfer overhead to exceed 30% in some scenarios.
  • After a modification to release the Python GIL promptly on the intra-node path, additional KV Transfer overhead fell from more than 30% to below 1%.
  • The Infra Agent improved a Decode Kernel's performance by 1.71 times by eliminating four repeated executions of the same group of normalization calculations.

woshipm用了一周后,来深入聊聊GLM-5.3

  • The latest Artificial Analysis ranking included three Chinese models in its top 10.
  • Kimi K3 and GLM-5.3 ranked fourth globally and first among Chinese models in the latest Artificial Analysis ranking.
  • GLM-5.3 has 744B total parameters and 40B active parameters.
  • GLM-5.3 scored 84.5% on CyberGym, compared with 83.8% for Mythos 5 and 83.6% for GPT-5.6 Sol.
  • Six models scored above 60 on the Artificial Analysis ranking, and GLM-5.3 had the smallest parameter count among them at 744B with 40B active parameters.
  • GLM-5.3 and Kimi K3 both scored 60 on Artificial Analysis, while GLM-5.3 has 744B parameters and Kimi K3 has 2.8T parameters.
  • DeepSeek-V4-Pro-0813 scored 53 on Artificial Analysis and has about 1.6T total parameters.
  • GLM-5.3 is available through Z.ai, the Zhipu Qingyan app, the BigModel Open Platform, and GLM Coding Plan.
  • GLM-5.3 has the same API pricing as GLM-5.2: 8 yuan per million input tokens, 2 yuan per million cached input tokens, and 28 yuan per million output tokens.

woshipm横评GLM-5.3、DeepSeek-v4-pro、K3,谁才是Coding新王?

  • The evaluation compared Kimi K3, GLM-5.3, and DeepSeek-v4-pro-0813 using the same task types, prompts, scoring standards, official APIs, and Claude Code.
  • The evaluation covered four coding tasks: frontend webpage development, 3D scene development, website development, and game development.
  • The evaluation assessed models on task completion and output quality.
  • For the frontend webpage task, the prompt asked for an HTML page explaining the differences between JPG/PNG and SVG, with a black technology-style background and advanced design.
  • For the 3D scene task, the prompt asked for a Three.js interactive webpage depicting 100,000 heavenly soldiers attacking Huaguo Mountain, with plot, viewpoint, and scenes matching Journey to the West.
  • DeepSeek-v4-pro delivered an unfinished 3D scene, and its page remained black and could not render normally after repeated debugging.
  • K3 and GLM-5.3 each produced their 3D scene output in one attempt.
  • For the website-development task, the models were asked to design a product promotional webpage for the open-source lengyi-title skill at https://github.com/woyin2024/lengyi-title.
  • GLM-5.3 is a text-only model trained after GLM-5.2 and does not have visual recognition capability.
  • For the game-development task, the models were asked to develop a 3D interactive web game based on the script file "源记-游戏.md".
  • Only K3's game version could be played from beginning to completion; GLM-5.3 had nonfunctional controls in the second level, while DeepSeek-v4-pro errored in the first level and unexpectedly jumped to the second level.
  • K3 implemented two narrative routes in the game: not reporting led to Ending One, "不复得路"; reporting led through leading a team to Ending Four, "黑色桃花", then the finale and true ending, "问津者".
  • For the most complex game case, DeepSeek-v4-pro took more than 10 minutes, GLM-5.3 took 43 minutes, and K3 took 67 minutes.