Anthropic Says China's Open-Weight GLM-5.3 Nearly Matches Its Own Model at Building Exploits

Anthropic Says China's Open-Weight GLM-5.3 Nearly Matches Its Own Model at Building Exploits

Anthropic Says China's Open-Weight GLM-5.3 Nearly Matches Its Own Model at Building Exploits

Anthropic's Frontier Red Team has published a report raising concerns about GLM-5.3, an open-weight model released by Chinese AI company Zhipu AI, also known as Z.ai. According to Anthropic, GLM-5.3 has reached a level of autonomous cyber-exploit development that approaches the capability of Claude Mythos Preview, an unreleased Anthropic model developed under what the company calls "Project Glasswing." Many observers note that the report frames this as evidence of advanced offensive AI capabilities spreading to models that, in Anthropic's assessment, lack adequate safeguards.

How GLM-5.3 Compares on Exploit-Development Benchmarks

Anthropic reports that on its internal ExploitBench evaluation, GLM-5.3 developed working end-to-end exploits in 50 of 410 attempts, compared to 56 of 410 for Claude Mythos Preview. On a separate Binary Exploitation benchmark, GLM-5.3 reportedly achieved full control-flow hijacks in 4% of trials, versus 6% for Mythos. By comparison, Anthropic says other open-weight models it tested, including Kimi K3 and DeepSeek V4.1 Flash, scored 0% on the same measure, which the report frames as making GLM-5.3 a notable outlier among open-weight systems. Anthropic also points to NIST's Center for AI Standards and Innovation, known as CAISI, which it says independently assessed GLM-5.3 as the most cyber-capable open-weight model currently available, offering a degree of third-party corroboration for its own findings.

Safeguards That Strip Away Easily

Anthropic states that GLM-5.3's stock refusal rates are comparable to those of its own models when evaluated normally. However, the report describes several techniques that reportedly bypass these safeguards at high rates. Deceptive roleplay-style prompts and a method called "prefilling" are said to succeed between 64% and 100% of the time, depending on the technique. The report places particular emphasis on "abliteration," a community-developed method for removing refusal behavior from open-weight models by modifying their internals. Anthropic reports that abliteration dropped GLM-5.3's refusal rate from above 90% to between 2% and 12% across three benchmarks, with what it describes as minimal loss of underlying capability. In one example cited in the report, an abliterated version of GLM-5.3-Flash reportedly produced a working ARM64 exploit chain bypassing pointer authentication hardening for a known vulnerability, requiring roughly eight hours of model computation and about 20 minutes of human attention.

Practical Barriers and Dual-Use Framing

Running an abliterated version of GLM-5.3 at meaningful scale reportedly requires substantial computing resources, with estimates citing around 306GB of VRAM and multi-GPU clusters such as eight H200 units. A recurring consumer concern raised alongside the report is whether this hardware barrier meaningfully limits real-world misuse, since such infrastructure is costly and not readily accessible to casual actors. Anthropic's report also includes dual-use framing, noting that the same underlying capabilities that could enable offensive exploit development may also support legitimate defensive security research, a point the company uses to balance its warnings about risk with acknowledgment of potential benefits.

Questions About the Messenger

Secondary coverage of the report, including analysis from Tom's Hardware, has raised questions about the circumstances surrounding its publication. Observers point out that Anthropic is a competing closed-weight AI lab reportedly preparing for an eventual public offering, and that it is reporting on the risks of a rival open-weight model developed by a competitor. Critics note that the benchmark results are self-reported by Anthropic and have not been independently replicated outside of CAISI's corroborating assessment. Some commentary has also questioned framing choices in the report, such as references to "Mythos-class" capabilities and descriptions of "weak safeguards," given that GLM-5.3's stock refusal rates are described in the same report as comparable to Anthropic's own models. A recurring observation is that the vulnerability to abliteration, while notable, is described by some as a less surprising finding given how commonly the technique is already applied to open-weight models generally. The broader context includes ongoing debate about the risks and benefits of open-weight AI distribution, as well as geopolitical framing that sometimes surrounds comparisons between Chinese and Western AI labs.

More A.I. articles · CuencaLife home