English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

GLM-5.3 Scores 84.5% on CyberGym: Zhipu's Open-Source Coding Model That Finds Other People's Bugs

Forum topic · 小凯 · 2026-08-17

Summary

On August 14, Zhipu released GLM-5.3, an open-weights coding model positioned as the strongest open model for coding to date. Without changing its base model—relying purely on post-training scaling—it improved Terminal-Bench from 4.6 to 28.3, reached 66.9 on DeepSWE, gained roughly 50% on the internal Z.ai Code Bench versus GLM-5.2, and will release open weights within two weeks. Most notably, GLM-5.3 scored 84.5% on the white-box cybersecurity benchmark CyberGym, matching Anthropic's closed-source flagship Mythos 5, and reportedly found 2,436 real-world vulnerabilities (1,097 rated medium-or-high severity) across 269 open-source projects including kernels, browsers, and DNS software. The article cautions that CyberGym uses white-box settings with ground truth, and zero-day discovery at scale remains unproven. Zhipu also launched an 'Open Source Shield' program offering free security audits for key open-source projects, framing safety capabilities as a public good.

The second half of the coding-model race is not about "writing more code"—it's about "finding code that others wrote wrong."

Zhipu's GLM-5.3, released August 14, makes this explicit: without swapping the base model, relying purely on post-training scaling, it jumped Terminal-Bench from 4.6 to 28.3, reached 66.9 on DeepSWE, improved roughly 50% on the internal Z.ai Code Bench versus GLM-5.2, and will be open-sourced within two weeks. One-line summary: this is currently the strongest open-source model for coding, and that "strongest" isn't just about writing code—cybersecurity capability emerged alongside it.

The numbers: from writing code to finding vulnerabilities

GLM-5.3 scored 84.5% on the white-box cybersecurity benchmark CyberGym, directly matching Anthropic's closed-source flagship Mythos 5. As of public release, the model had cumulatively discovered 2,436 vulnerabilities in real environments, 1,097 of which were rated medium or high severity, covering 269 open-source projects including kernels, browsers, and DNS. These are not results a "code completion" side project can produce—finding them requires understanding the context of large codebases while actively reasoning about attack surfaces, skills that traditional coding benchmarks don't test at all.

More interesting is how Zhipu wrote the fusion of "coding model" and "security model" into its product narrative. Alongside the release, Zhipu launched an "Open Source Shield" program, announcing free security audits for key open-source projects, treating frontier security capabilities as a public good that "should be open rather than a closed-source privilege." This is a new positioning: model companies are no longer just writing tools for developers, but turning models into the "immune system" of the open-source ecosystem.

Limitations: CyberGym is not the real world

A dose of cold water for the numbers. CyberGym is a white-box benchmark with ground truth; validating real vulnerabilities still depends on CVE numbers, manual reproduction, and disclosure processes. There is no public evidence yet that the model can maintain similar scores on long-chain tasks like "zero-day vulnerability discovery." In other words, GLM-5.3 is very good at "finding holes where holes are known to exist," but "digging out new holes where nobody thinks there are holes" remains an engineering challenge of a different magnitude.

Also, the 84.5% figure should be read as "a natural byproduct that emerges once coding capability reaches a certain level," not as GLM-5.3's design goal. It is not a model specifically trained for security auditing—its code understanding is simply deep enough that understanding vulnerabilities comes along for the ride.

What it changes

The "next racing point" in the coding-model track has quietly been swapped. Over the past six months, everyone competed on "can it run SWE-Bench, HumanEval, LiveCodeBench." Starting with GLM-5.3, the boundaries of this track have been pushed out: whether a model can write code, audit code, and find vulnerabilities in other people's code will become the new capability dividing line.

On a bigger scale, this means open-source models are shifting from "performance catching-up" to "capability spillover." The gap between open and closed models is no longer in coding itself, but in the byproducts beyond coding—security, operations, documentation, observability. GLM-5.3 is the first sample line showing that "a coding agent's moat lies not in writing, but in reviewing."

Over the next 12 months, three validation points are worth watching: first, whether DeepSeek-V5, Qwen3.9, and Kimi K4 also put CyberGym-style benchmarks into their release materials; second, whether the "Open Source Shield" program can commercialize GLM-5.3's security-audit capability into a SaaS package within 6 months; third, whether a "double-crown champion" of coding + security capability becomes the default release posture for the next generation of models.

There are plenty of models that write code. An open-source coding model that audits other people's code—this is the first time it has been formally put on the table.

Tags

#glm-5#zhipu#open-source-models#coding-llm#cybersecurity#vulnerability-discovery#cybergym#ai-benchmarks

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178633571