Signal

Anthropic says GLM-5.3 can autonomously build end-to-end cyber exploits, like Claude Mythos Preview, but was released without robust safeguards against misuse

First reported by Anthropic ·

The signal ●●●○ Compiled by AI from Anthropic, Techmeme and Simon Willison's Weblog
Why you might care

Openly available AI models can now autonomously build cyber exploits, lowering the barrier to entry for attackers.

What happened

Anthropic researchers have analyzed GLM-5.3, a new AI model from Zhipu AI (Z.ai), and found it possesses advanced capabilities for autonomously building end-to-end cyber exploits, similar to Anthropic's own Claude Mythos Preview. Unlike Claude Mythos Preview, which was released under strict controls to vetted users for defensive purposes, GLM-5.3 has been released as an open-weight model with what Anthropic describes as "meaningful safeguards to limit misuse" that can be bypassed or removed with simple techniques. Anthropic's testing showed attackers can bypass GLM-5.3's safeguards between 64% and 100% of the time, and that abliterating the model (removing its refusals) is feasible with significant but attainable computational resources. NIST's Center for AI Standards and Innovation (CAISI) also assessed GLM-5.3, deeming it the "most cyber-capable open-weight model released to date" and lagging behind U.S. frontier models by approximately four months in aggregate cyber benchmarks.

What it means

The release of GLM-5.3 without robust safeguards marks a significant shift, democratizing the creation of sophisticated cyber exploits. While Anthropic's Claude Mythos Preview was intentionally kept from the public to aid defenders, GLM-5.3's open-weight nature means anyone can download and modify it, potentially enabling a surge in novel and impactful cyberattacks. This development compels a re-evaluation of AI safety protocols, particularly for models with demonstrable offensive capabilities that are not tightly controlled.

The ease with which GLM-5.3's safeguards can be bypassed, even through techniques like "abliteration," suggests that current open-weight models may not offer sufficient protection against misuse. This has direct implications for cybersecurity, as defenders must now contend with AI-generated exploits that could be developed and deployed with unprecedented speed and reduced cost. The race is on for both AI developers to implement more resilient safeguards and for security teams to adapt their threat detection and mitigation strategies.

AI-written summary. May contain errors.