On August 14, Zhipu (智谱) released GLM-5.3. The base model is identical to GLM-5.2 — still ~743B parameters, with no architecture changes and no added parameters. Every improvement comes from post-training scaling. Coding capability improved 50% over the previous generation; Terminal Bench 3.0 jumped from 4.6 to 28.3 (open-source #1, a 6x increase), and DeepSWE v1.1 rose from 46.2 to 66.9 (open-source #1). On Z.ai Code Bench, GLM-5.3 High reached 31.4% accuracy, surpassing Claude Opus 4.8 Max's 29.5%, while consuming only 50K tokens per task — Opus 4.8 needs 120K.
Hard Evidence for Same-Base Improvement
Zhipu made the accounting clear: GLM-5.3 uses the same 743B base as GLM-5.2, but the post-training stage moved to a new generation Slime framework (the framework itself is open-sourced), introducing reinforcement learning scheduling such as IndexShare and SAO, and trained the model for extended periods in long-horizon task environments.
| Benchmark | GLM-5.2 | GLM-5.3 | Gain | |---|---|---|---| | Terminal Bench 3.0 | 4.6 | 28.3 | 6x | | DeepSWE v1.1 | 46.2 | 66.9 | +44.8% | | Agents' Last Exam | 23.8 | 28.5 | +19.7% | | CyberGym | 77.2 | 84.5 | +7.3 | | AutomationBench | 26.2 | 48.2 | +22 | | HLE w/Tools | 54.7 | 62.5 | +7.8 | | GDPval-AA v2 | 1508 | 1769 | +261 |
Five of six benchmarks rank open-source #1, one ranks #2. Terminal Bench 3.0 is the gold standard for whether a model can independently complete complex tasks in a real terminal environment; the jump from 4.6 to 28.3 means GLM-5.2 could barely use a terminal, while GLM-5.3 can handle work approaching an engineer's day.
Coding Efficiency: Same Work at Half the Tokens
Z.ai Code Bench runs on a Claude Code 2.1.207 evaluation harness, and the result is counterintuitive — GLM-5.3 High completes tasks with 50K tokens that Opus 4.8 needs 120K tokens for (with accuracy 1.9 percentage points higher). That works out to 2.4x token efficiency.
Practical meaning: in the API-billing era, tokens are money. GLM-5.3's 50% coding gain translates into inference-stage token savings rather than higher user bills. Zhipu explicitly highlighted "50K vs 120K" to directly target Opus 4.8 customers.
Emergent Cybersecurity: 2,404 Vulnerabilities, Some 40 Years Old
Post-training scaling unexpectedly activated cybersecurity capability. Zhipu worked with domestic security teams on two weeks of intensive testing before release; GLM-5.3 found 2,404 potential vulnerabilities across 269 repositories, of which 1,088 are medium-to-high severity. The oldest had lain dormant for over 40 years (an architectural flaw in the DNS protocol itself, where a small number of malformed requests can amplify traffic 80,000x).
On CyberGym (white-box code review), GLM-5.3 scored 84.5%, beating Mythos 5 (83.8%) and GPT-5.6 Sol (83.6%) for first place overall. On ExploitBench it scored 54.4, more than double GLM-5.2's 24.4.
Zhipu launched a cvd.z.ai security disclosure ledger, publicly registering vulnerabilities by CVE status, severity, and latency age; 2,383 remain in coordinated disclosure. This "proactive disclosure + delayed open-weight" cadence is rare restraint among domestic models in this wave.
Commercial Rollout: Two-Week Lock, Then Open Weights
GLM-5.3 is now fully available on ZCode, AutoClaw, and GLM Coding Plan, with early access on Trae, Qoder, JoyCode, OpenCode, CodeBuddy, WorkBuddy, and other developer platforms. Full model weights will be open-sourced two weeks later — a "hardening window" for security teams, and the first time a leading domestic model has standardized a release → harden → open-source pipeline.
Two signals to watch over the next two weeks: 1) the number of community fine-tuned derivatives within 24 hours of weight release; 2) whether GLM Coding Plan subscription growth matches Claude Code (whose weekly retention is ~70%).
Sources: Zhipu official announcement, Singularity.kiwi vulnerability disclosure analysis, GitHub AI Daily 8-15.