2026-08-13

GLM-5.3: Same Model, Two Weeks of Post-Training, and a Capability Z.ai Says Surprised Them

AISecurity🌍 Asia

Z.ai released GLM-5.3 today, and the framing is unusually clean as a natural experiment: "it uses the same base model as GLM-5.2 — every gain comes from post-training." No new pretraining, no architecture change — the entire jump comes from a month of scaling the same RL stack (IndexShare, SAO, slime) across more environments and more compute. That makes GLM-5.3 a rare clean read on how much post-training scaling alone is worth right now, and the answer, on this evidence, is: a lot.

Three outright wins against the closed frontier

Z.ai's own comparison table — GLM-5.3 against GLM-5.2, Kimi K3, DeepSeek-V4-Pro-0813, Qwen3.8-Max, Claude Opus 4.8, Claude Fable 5, and GPT-5.6 Sol — shows GLM-5.3 leading outright on three benchmarks, beating every closed model shown: CyberGym (84.5, ahead of Fable 5's 83.8), AutomationBench (48.2, ahead of Kimi K3's 46.7), and GDPval-AA v2 (1769, ahead of Fable 5's 1743 and Qwen3.8-Max's 1739). That's a different result from what Qwen3.8-Max and DeepSeek-V4-Pro delivered this month — both closed most of the gap to frontier closed models without crossing it. GLM-5.3 crosses it, on three specific axes, while still trailing Fable 5 and GPT-5.6 Sol on most of the coding-agent cluster (DeepSWE, NL2Repo, FrontierSWE, Toolathlon Verified).

The efficiency story compounds this: on Z.ai's own Code Bench, at High effort GLM-5.3 hits 31.4% accuracy at ~50K output tokens, ahead of Claude Opus 4.8's 29.5% at 120K — better and less than half the tokens. It's still behind Claude Fable 5 (39.5% at Max effort), but beating a Claude-tier model on both axes at once, not just accuracy, is a genuinely sharper result than a benchmark win alone.

The cyber story is the one to slow down for

Z.ai added vulnerability-discovery data and environments to post-training expecting a modest capability bump. Their own words: "What surprised us was how quickly the capability continued to develop as training scaled." GLM-5.3 didn't just get better at spotting isolated flaws — it started reasoning across multiple stages of exploitation, forming coherent multi-step exploitation plans. On CyberGym (white-box vulnerability discovery), it's SOTA at 84.5%. On ExploitBench (deeper exploitation reasoning), it more than doubles GLM-5.2, 54.4% against 24.4% — though Fable 5 and GPT-5.6 Sol remain well ahead there, at 78.0% and 76.5%. Z.ai's own read: "the further up the exploitation chain a benchmark sits, the larger the gain from GLM-5.2 — and also the wider the remaining gap to the closed frontier. Capability is growing fastest exactly where we are furthest behind." That's a lab watching a dual-use capability accelerate in training and saying so plainly, rather than discovering it after deployment the way OpenAI and Hugging Face did in July.

Then it gets real: Z.ai ran GLM-5.3 against actual production codebases with Chinese security teams, and after expert review and deduplication, it found 2,436 vulnerabilities across 269 real projects — 1,097 of them medium-to-high severity — in system kernels, browser engines, operating systems, and network protocols. The oldest bug dates to 1981; the average finding sat undiscovered for 26.6 years. Z.ai built a public Security Disclosure Ledger tracking all of it: 53 disclosed publicly so far, 2,383 still under embargo. This is not a benchmark claim — it's a verifiable, ongoing public record of an AI model finding real, exploitable flaws in software the world actually runs, at a scale and hit rate no benchmark table captures.

Which makes the release plan the most consequential sentence in the announcement: open weights ship two weeks after launch, "once safety evaluation and hardening are complete." That's the gated-release logic we've tracked all month — OpenAI, Anthropic, and Google all shipped access-restricted "cyber" model tiers rather than open ones — except here it's an open-weights-by-default lab building in an explicit delay for the first time, on the capability that most needs one. The UK AI Safety Institute's estimate that open-weight models trail the frontier on cyber capability by only 4–7 months reads differently now than it did in July: GLM-5.3 didn't trail on discovery, it led, and the two-week countdown to open weights is the actual test of whether "safety evaluation and hardening" changes what ships or just delays it.

Two things converging with this week's other launches

Mandatory reasoning. GLM-5.3 drops support for disabled thinking entirely — three effort levels (low, high, max), no off switch, existing integrations using thinking.type: "disabled" will hard-fail until migrated. That's the same shape Qwen3.8-Max and DeepSeek-V4-Pro shipped this month — three separate labs now treating always-on reasoning as the default for their flagship models, not a toggle.

Peak/off-peak pricing. The GLM Coding Plan now runs a points-based quota with calls outside 14:00–18:00 (UTC+8) weekdays billed at 50% off. One day after we flagged DeepSeek's time-of-day surge pricing as a structural first for the industry and asked whether a second lab would adopt it within the quarter — a second lab adopted it within a day. That's a fast enough turnaround to suggest this wasn't independently invented twice; it's becoming the obvious move once demand outstrips capacity at these prices.

What to expect next

  • Watch the two-week open-weights window closely. Whether the released weights match the demonstrated cyber capability, ship with mitigations, or slip past two weeks will say more about how seriously "safety hardening" is being taken than the announcement does.
  • Watch the Security Disclosure Ledger. A live, public, ongoing record of AI-discovered vulnerabilities moving through disclosure is a genuinely new kind of artifact — worth checking back on as more of the 2,383 embargoed findings clear.
  • Watch whether GLM-5.3's three outright wins get independently reproduced. Beating Claude and GPT-5.6 Sol outright, even on three specific benchmarks, is the kind of claim worth checking against a source outside the vendor before treating as settled.

References: Z.ai — GLM-5.3: Frontier Coding with Emergent Cyber Capabilities · related coverage: DeepSeek-V4-Pro · The Cyber Model Trend · When the Eval Escaped · Qwen3.8-Max's open weights land · Gemini 3.7 Flash · Frontier Arcade: trends & predictions