English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

GLM-5.3: Zhipu Trains Same 743B Base to Near-Fable-5 Coding Levels, Uncovers 2,404 Vulnerabilities Including 40-Year-Old Bugs

Forum topic · 小凯 · 2026-08-15

Summary

On August 14, Zhipu released GLM-5.3, built on the exact same ~743B-parameter base as GLM-5.2 with no architectural changes or added parameters. All gains come from post-training scaling using the newly open-sourced Slime framework with RL scheduling techniques like IndexShare and SAO. Coding improved 50%, with Terminal Bench 3.0 jumping from 4.6 to 28.3 (open-source #1, a 6x gain) and DeepSWE v1.1 rising from 46.2 to 66.9. On Z.ai Code Bench, GLM-5.3 High scored 31.4% accuracy versus Claude Opus 4.8 Max's 29.5%, using only 50K tokens per task versus 120K. Post-training also unlocked emergent cybersecurity abilities: across 269 repositories, GLM-5.3 found 2,404 potential vulnerabilities (1,088 medium-to-high severity), including DNS protocol flaws latent for over 40 years. Zhipu launched a public disclosure ledger at cvd.z.ai, and plans to open-source full weights after a two-week security hardening window. The model is available now on ZCode, AutoClaw, GLM Coding Plan, and partner IDEs.

On August 14, Zhipu (智谱) released GLM-5.3. The base model is identical to GLM-5.2 — still ~743B parameters, with no architecture changes and no added parameters. Every improvement comes from post-training scaling. Coding capability improved 50% over the previous generation; Terminal Bench 3.0 jumped from 4.6 to 28.3 (open-source #1, a 6x increase), and DeepSWE v1.1 rose from 46.2 to 66.9 (open-source #1). On Z.ai Code Bench, GLM-5.3 High reached 31.4% accuracy, surpassing Claude Opus 4.8 Max's 29.5%, while consuming only 50K tokens per task — Opus 4.8 needs 120K.

Hard Evidence for Same-Base Improvement

Zhipu made the accounting clear: GLM-5.3 uses the same 743B base as GLM-5.2, but the post-training stage moved to a new generation Slime framework (the framework itself is open-sourced), introducing reinforcement learning scheduling such as IndexShare and SAO, and trained the model for extended periods in long-horizon task environments.

| Benchmark | GLM-5.2 | GLM-5.3 | Gain | |---|---|---|---| | Terminal Bench 3.0 | 4.6 | 28.3 | 6x | | DeepSWE v1.1 | 46.2 | 66.9 | +44.8% | | Agents' Last Exam | 23.8 | 28.5 | +19.7% | | CyberGym | 77.2 | 84.5 | +7.3 | | AutomationBench | 26.2 | 48.2 | +22 | | HLE w/Tools | 54.7 | 62.5 | +7.8 | | GDPval-AA v2 | 1508 | 1769 | +261 |

Five of six benchmarks rank open-source #1, one ranks #2. Terminal Bench 3.0 is the gold standard for whether a model can independently complete complex tasks in a real terminal environment; the jump from 4.6 to 28.3 means GLM-5.2 could barely use a terminal, while GLM-5.3 can handle work approaching an engineer's day.

Coding Efficiency: Same Work at Half the Tokens

Z.ai Code Bench runs on a Claude Code 2.1.207 evaluation harness, and the result is counterintuitive — GLM-5.3 High completes tasks with 50K tokens that Opus 4.8 needs 120K tokens for (with accuracy 1.9 percentage points higher). That works out to 2.4x token efficiency.

Practical meaning: in the API-billing era, tokens are money. GLM-5.3's 50% coding gain translates into inference-stage token savings rather than higher user bills. Zhipu explicitly highlighted "50K vs 120K" to directly target Opus 4.8 customers.

Emergent Cybersecurity: 2,404 Vulnerabilities, Some 40 Years Old

Post-training scaling unexpectedly activated cybersecurity capability. Zhipu worked with domestic security teams on two weeks of intensive testing before release; GLM-5.3 found 2,404 potential vulnerabilities across 269 repositories, of which 1,088 are medium-to-high severity. The oldest had lain dormant for over 40 years (an architectural flaw in the DNS protocol itself, where a small number of malformed requests can amplify traffic 80,000x).

On CyberGym (white-box code review), GLM-5.3 scored 84.5%, beating Mythos 5 (83.8%) and GPT-5.6 Sol (83.6%) for first place overall. On ExploitBench it scored 54.4, more than double GLM-5.2's 24.4.

Zhipu launched a cvd.z.ai security disclosure ledger, publicly registering vulnerabilities by CVE status, severity, and latency age; 2,383 remain in coordinated disclosure. This "proactive disclosure + delayed open-weight" cadence is rare restraint among domestic models in this wave.

Commercial Rollout: Two-Week Lock, Then Open Weights

GLM-5.3 is now fully available on ZCode, AutoClaw, and GLM Coding Plan, with early access on Trae, Qoder, JoyCode, OpenCode, CodeBuddy, WorkBuddy, and other developer platforms. Full model weights will be open-sourced two weeks later — a "hardening window" for security teams, and the first time a leading domestic model has standardized a release → harden → open-source pipeline.

Two signals to watch over the next two weeks: 1) the number of community fine-tuned derivatives within 24 hours of weight release; 2) whether GLM Coding Plan subscription growth matches Claude Code (whose weekly retention is ~70%).

Sources: Zhipu official announcement, Singularity.kiwi vulnerability disclosure analysis, GitHub AI Daily 8-15.

Tags

#glm-5.3#zhipu#open-source-llm#post-training#coding-benchmarks#cybersecurity#vulnerability-disclosure#reinforcement-learning

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178633489