Technology

A New AI Model from China Nearly Matches Claude on Coding — at a Tenth of the Price

Martin HollowayPublished 5w ago5 min readBased on 7 sources
Reading level
A New AI Model from China Nearly Matches Claude on Coding — at a Tenth of the Price
Photo by Brett Sayles on Pexels

On August 26, 2026, a Chinese company called Z.ai (also known as Zhipu AI) released a new AI model called GLM-5.3-Flash. It is the first in its GLM-5 series that can work with multiple types of input — text, images, and more — from the ground up, rather than handling them as separate steps. The model is priced at roughly one-tenth the cost of its predecessor GLM-5.2, yet it outperforms that older model across benchmarks and real-world workloads. Z.ai Blog

To understand the size of this model: it has 320 billion parameters in total (parameters are the internal settings or weights that determine what the model knows), but only 18 billion of them are "active" during any single request. Think of it like a large library where only the relevant shelves get consulted for each question, rather than searching every shelf every time. This approach, called sparse activation, keeps costs down while maintaining quality.

The model uses a combination of attention mechanisms — the techniques AI models use to decide which parts of the input to focus on. One is sparse attention, which skips irrelevant portions of the input, and the other is linear attention, which handles long stretches of text more efficiently. Together they reduce the cost of serving very long conversations while keeping quality intact. A feature called IndexPool compresses four internal tracking vectors into one, cutting the time and memory the model needs when processing up to one million tokens of context at once. Compared to the larger GLM-5.3, the Flash variant cuts attention computation by 3x and its temporary memory usage by 4.4x.

The parameter economics are notable. Against the GLM-4.5 series, GLM-5.3-Flash holds a similar total parameter count (320B vs 355B) but nearly halves activated parameters (18B vs 32B) and slashes the layer count from 92 to 45. Among GLM-5.3-Flash, GLM-5.3, DeepSeek-V4-Flash, and Kimi-K3, GLM-5.3-Flash achieves the lowest attention compute per head per layer, though its temporary memory usage remains slightly larger than Kimi-K3 and DeepSeek-V4-Flash.

The model was trained on a dataset of 30 trillion tokens drawn from multiple input types (text, images, and other data combined). The base model, GLM-5.3-Flash-Base, outperforms GLM-4.5-Base overall and remains competitive with GLM-5-Base across most benchmarks. At 18B activated / 320B total, it sits between DeepSeek-V4-Flash-Base (13B/284B) and GLM-4.5-Base (32B/355B), with GLM-5-Base at 40B/744B. Z.ai Blog

Before the official announcement, Z.ai quietly tested the model under the codename "ox-alpha" on two popular AI platforms — OpenCode and OpenRouter. Under that anonymous name, it became the most popular model of the week. All of that traffic was served on Chinese-made AI chips. Z.ai Blog

On benchmark tests, GLM-5.3-Flash approaches Claude Opus 4.8 — one of the top AI models for coding — on coding and agentic tasks (tasks where the AI acts semi-autonomously to complete multi-step work). On the DeepSWE v1.1 benchmark it scored 63.4 versus 46.2 for GLM-5.2; on AutomationBench it scored 48.8 versus 26.2. On Z.ai Code Bench v1.0 (run on Claude Code 2.1.207), it outperforms GLM-5.2 at every effort level, and at maximum effort scores 29.0 versus 29.5 for Claude Opus 4.8. On the Artificial Analysis Intelligence Index v4.1.1, GLM-5.3-Flash scores 57 at $0.045 per task (discounted), placing it on the Pareto frontier of that index — meaning it offers more capability per dollar than any model previously tracked there. Z.ai Blog

Z.ai offers a GLM Coding Plan that supports GLM-5.3, GLM-5.2, and GLM-5-Turbo for AI-assisted coding in tools including Claude Code, Kilo Code, Cline, OpenCode, and Clawdbot/OpenClaw. Z.ai

The GLM-5 series was designed by Z.ai for complex systems engineering and long-horizon agentic tasks, with GLM-5.2 introduced in June 2026 as the flagship for long-horizon capability. Z.ai Blog The release notes also indicate GLM-5.3 delivers a 50% gain over GLM-5.2 on Z.ai Code Bench, reaching state-of-the-art in coding capabilities. Z.ai Docs

The broader context here is about cost and infrastructure. A model approaching Claude Opus 4.8 on coding benchmarks at $0.045 per task, while being served entirely on domestic Chinese silicon, signals something concrete about the trajectory of inference economics. The ox-alpha blind test is worth pausing on: developers chose the model on merit, not brand, and the infrastructure question was answered only after the fact. That the serving hardware was Chinese AI chips, disclosed post-hoc, adds a layer of significance for anyone tracking the semiconductor supply chain and its relationship to model deployment at scale.

The architecture choices tell their own story. Halving activated parameters and layer count relative to GLM-4.5, while holding total parameters roughly constant, is a bet that sparser activation and a shallower effective depth can maintain quality when paired with hybrid attention and techniques like IndexPool. The memory footprint tradeoff against Kimi-K3 and DeepSeek-V4-Flash suggests the design prioritized attention compute reduction over memory footprint — a reasonable call when serving cost is dominated by compute rather than memory bandwidth at the target price point.

For practitioners, the practical question is whether GLM-5.3-Flash's coding and agentic benchmark performance translates into reliable real-world utility in the agentic coding workflows where Claude has built a strong reputation. The Z.ai Code Bench numbers at max effort (29.0 vs 29.5 for Claude Opus 4.8) are close enough to warrant direct evaluation in production toolchains, and the GLM Coding Plan's integration with Claude Code, Cline, and OpenCode lowers the barrier to that testing. The fact that the model was already serving real traffic on OpenRouter under a codename means some of that real-world signal already exists.