deniz.in

Markets

Weather

Loading weather

· via Hacker News – Front Page (hnrss.org)

Anthropic finds Zhipu's GLM-5.3 builds autonomous end-to-end exploits with weak safeguards

Anthropic says Zhipu AI's GLM-5.3 can build end-to-end cyber exploits at close to Claude Mythos Preview's level while being freely downloadable, and its safeguards were bypassed 64–100% of the time.

Anthropic finds Zhipu's GLM-5.3 builds autonomous end-to-end exploits with weak safeguards

Anthropic has published an analysis of GLM-5.3, the newest model from Zhipu AI (known outside of China as Z.ai), and its conclusion is blunt: the model can autonomously develop working cyber exploits from discovery to execution, roughly matching a capability Anthropic deliberately kept under wraps, while its weights are downloadable by anyone. The post, written by five Anthropic researchers — Andrew Fasano, Marius Fleischer, Cole McFaul, Robert Xiao and Tripp Gallagher — surfaced on the Hacker News front page on September 29.

The capability Anthropic held back

Five months earlier, Anthropic announced Claude Mythos Preview, which it describes as the first AI model able to build sophisticated exploits across an entire attack chain without step-by-step human help. Anticipating that this ability would spread to other models, the company released it only through Project Glasswing, a channel for vetted defenders, who used it to surface more than 10,000 vulnerabilities in critical software before malicious actors had comparable tools. In the new post, Anthropic says models of that class have now arrived, starting with GLM-5.3.

How GLM-5.3 scored

All evaluations ran in sandboxed environments against offline targets set up for testing. On ExploitBench, which measures exploitation of known bugs in the V8 engine behind Google Chrome, GLM-5.3 completed a working exploit in 50 of 410 attempts, close to Claude Mythos Preview's 56 of 410. On Anthropic's internal binary exploitation benchmark, drawn from Google OSS-Fuzz projects and awarding full credit only for a complete control-flow hijack, GLM-5.3 solved 4% of 100 randomly selected tasks versus 6% for Mythos Preview. The gap that matters, according to Anthropic, is the baseline: earlier models such as Claude Opus 4.6 and GLM-5.2 succeeded on none of them. The comparison figures also include the latest open-weight releases from Moonshot AI (Kimi K3) and DeepSeek (V4.1-Flash).

What researchers achieved with the model

In the first human-driven session, a researcher pointed GLM-5.3 at a sandboxed Linux build of a popular web browser. Over the course of a day, with limited human attention, the model found several previously unknown vulnerabilities in the browser's JavaScript engine and chained them into a malicious page that reads arbitrary files from a visitor's computer — demonstrated by pulling a user's SSH private key. Anthropic has disclosed these bugs to the maintainer and believes users on other platforms could be affected, though exploitation there may be more complex. The same session turned up exploitable weaknesses in wireless and graphics drivers and in network-facing device software, with disclosures pending Anthropic's review.

A second session used GLM-5.3-Flash, a smaller sibling model. Given only public details of CVE-2026-11645, a recently disclosed Chrome flaw, plus one other known bug, the model combined the two into a reliable exploit chain against an ARM64 target that bypassed pointer-authentication hardening. The total effort came to 20 minutes of human attention, eight hours of model work, and $20.40 at Zhipu's API prices.

Safeguards that give way

GLM-5.3 does refuse clearly harmful requests, but in Anthropic's simulated tests, simple techniques bypassed or removed these protections 64% to 100% of the time. The same attacks did not succeed against safeguarded Claude models. The post identifies a standard refusal-reduction technique as its most effective method, though the published text cuts off mid-description.

An independent read from NIST

On September 17, NIST's Center for AI Standards and Innovation published its own assessment, calling GLM-5.3 "the most cyber-capable open-weight model released to date" and estimating that it trails the US frontier by about four months on an aggregate of its cyber benchmarks. Anthropic says its capability findings broadly agree, with one crucial caveat: the US models in that comparison were tested with safeguards disabled, and the frontier group includes models released only to vetted users. Attackers cannot readily reach those versions — but anyone can download GLM-5.3.

Why it matters

Anthropic frames GLM-5.3 as evidence that frontier offensive-cyber capability does not stay contained: a model matching a system Anthropic restricted to vetted defenders is now in open circulation with guardrails that can be stripped off. The practical shift is economic — turning a freshly patched browser bug into a working attack cost tens of dollars and a few hours, not a funded team's weeks. The same capability helps defenders find and fix flaws first, an upside Anthropic explicitly notes, but the model is equally available to both sides. Expect sharper debate over open-weight release practices, and heavier load on coordinated disclosure pipelines, as more labs cross this threshold.

  • #ai-security
  • #cybersecurity
  • #open-weights
  • #llm
  • #vulnerability-research

Related posts