Skip to main content
Aggregate CSO Online 网络安全 17 Aug 2026 - 20:32

Zhipu says new coding AI developed advanced cyber skills faster than expected

RSS 官方收录 · 可信分层展示

关键摘要

Chinese AI developer Zhipu has launched GLM-5.…

  • 3, a new coding-focused AI model that the company says has developed u…
  • Zhipu’s own testing places GLM-5.
  • 3 slightly ahead of Anthropic’s Mythos 5 and OpenAI’s GPT-5.

摘要引擎:抽取

正文提要

Chinese AI developer Zhipu has launched GLM-5.3, a new coding-focused AI model that the company says has developed unexpectedly strong cybersecurity capabilities, putting it close to global leading models in vulnerability discovery while remaining behind them on deeper exploitation tasks.

Zhipu’s own testing places GLM-5.3 slightly ahead of Anthropic’s Mythos 5 and OpenAI’s GPT-5.6 Sol on CyberGym, a benchmark that tests vulnerability identification and validation. GLM-5.3 scored 84.5%, compared with 83.8% for Mythos 5 and 83.6% for GPT-5.6 Sol. But the model trails both competitors by a much wider margin on ExploitBench, where it scored 54.4%, compared with 78% for Mythos 5 and 76.5% for GPT-5.6 Sol.

“GLM-5.3 is the most capable open-weights model for coding, with a 50% improvement over GLM-5.2 on our in-house Z.ai Code Bench,” Zhipu said in a statement. “As we scaled post-training, cyber capability developed faster than we expected.” The company said GLM-5.3 moved beyond identifying isolated vulnerabilities to “forming coherent plans for complete exploitation chains.”

The company also claimed its latest model has shown improvement over Zhipu’s previous GLM-5.2. Its ExploitBench score more than doubled from 24.4%, while on ExploitGym, GLM-5.3 completed 105 exploitation tasks within two hours and 130 within six hours, compared with 29 and 39 respectively for GLM-5.2, according to the statement.

Zhipu attributes the gains to post-training, including reinforcement learning across increasingly complex task environments.

The progression reflects a broader issue emerging as coding models become more capable, said Neil Shah, VP for research and partner at Counterpoint Research.

“We are reaching a stage where if we teach an AI to be a brilliant software engineer, you’re accidentally teaching it how to be a good hacker, too,” Shah said. “The exact same reasoning an AI uses to test code and fix bugs is what an attacker uses to find a weak spot and break through it.”

He said offensive cyber capability is becoming an inherent capability of next-generation coding AI, making controls around such systems an increasingly important issue.

Thousands of vulnerabilities found

Zhipu said it has also been working with security teams in China to test its models against real-world codebases.

“After expert review, screening, and deduplication, the model identified 2,436 vulnerabilities across 269 projects, including 1,097 medium-to-high severity issues,” the statement added.

The findings cover system kernels, operating systems, browser engines, open-source infrastructure, Web applications and network protocols, Zhipu said.

Zhipu’s security disclosure ledger lists 107 critical and 990 high-severity findings. The company said 53 findings have been publicly disclosed and 2,383 remain under embargo. The oldest vulnerability identified dates to 1981, while vulnerabilities in the dataset had remained in code for an average of 26.6 years before discovery.

The company did not disclose how many of the 2,436 findings were previously unknown vulnerabilities or how many were independently reproduced. It said the findings are being tracked through its Z.ai Security Disclosure Ledger as they move through the disclosure process.

Shah described the capability as a double-edged development for security teams.

“These AI tools can audit systems and fix bugs faster,” he said. “But once an AI model’s weights are released freely to the public, any built-in safety guardrails can be stripped away without any cognizance or control.”

Same base model, scaled post-training

Zhipu attributes GLM-5.3’s gains to scaling post-training rather than developing a new base model.

The company expanded its training environments to simulate longer and more realistic units of professional work. In one example, the model is given access to compute clusters, storage systems, internal documentation, codebases and experiment results and must diagnose a bottleneck, implement an optimization, run experiments and deliver a measurable improvement while maintaining correctness.

Zhipu also added vulnerability-discovery data and environments to the training mix.

The company reported a 50% improvement over GLM-5.2 on its internal Z.ai Code Bench, alongside gains on public coding and agent benchmarks.

Shah said the connection between coding and offensive security is becoming harder to separate as these models improve.

“The exact same reasoning an AI uses to test code and fix bugs is what an attacker uses to find a weak spot and break through it,” he said.

Open-weight release raises the stakes

Zhipu plans to release GLM-5.3’s model weights about two weeks after launch, following safety evaluation and hardening.

The company is preparing to make an open-weight model available that it says has demonstrated capabilities ranging from vulnerability discovery to increasingly sophisticated exploitation reasoning.

Zhipu has not said in the announcement what additional safeguards will accompany the open-weight release beyond its planned safety evaluation and hardening.

For Shah, the issue is the speed at which vulnerabilities could potentially move from discovery to exploitation once such capabilities are widely available.

“If these AI-driven tools can discover thousands of unpatched flaws in real-world systems and anyone can download that capability, the response window shrinks to near zero,” he said.

He said defending against attacks operating at machine speed would require controls built into the development and deployment of AI models and autonomous agents.

The article originally appeared on InfoWorld.

打开官方原文 站点原文页 可信分区 本信源更多 今日简报 分享图 RSS 稍后再看列表