Published
- 5 min read
When Anthropic Talks About Its GLM Rival: The Cyber AI Line Has Moved

Anthropic has published a security assessment of GLM-5.3, an open-weight model developed by Zhipu AI (Z.ai), and the subtext is hard to miss: a model outside Anthropic’s own ecosystem may now approach its frontier systems at certain exploit-development tasks while being much easier to obtain and, according to Anthropic’s tests, much easier to misuse.
That’s a deliberately provocative comparison. But the important story isn’t that one lab is criticizing a competitor. It’s that advanced cyber capabilities are spreading beyond restricted-access models.

What Anthropic says GLM-5.3 can do
Anthropic tested GLM-5.3 in isolated, sandboxed evaluations and reports that it:
- built end-to-end exploits in 50 of 410 ExploitBench attempts, close to Claude Mythos Preview’s 56 of 410;
- achieved full control-flow hijacks in 4% of a 100-task internal binary-exploitation test, compared with 6% for Mythos Preview;
- helped researchers find and chain previously unknown browser-engine vulnerabilities in a controlled session; and
- helped build an exploit chain for a known Chrome vulnerability, according to Anthropic, with about 20 minutes of human attention and eight hours of model work.
These results are not a guarantee that the model will reliably exploit arbitrary targets. They come from particular benchmarks and research setups. Anthropic also says its hands-on evaluations were conducted in sandboxes and that the unknown vulnerabilities were disclosed to maintainers or were under review.
The headline: capability is spreading, safeguards are not keeping pace
Anthropic’s sharpest finding concerns safeguards. In its simulated tests, the lab reports that GLM-5.3 engaged with harmful cyber requests 64% of the time with a deceptive red-team framing, 92% with prefilling, and 100% after the model was modified to reduce refusals. Anthropic says its tested safeguarded Claude models did not engage under the corresponding applicable conditions.
Those numbers need context: the experiments were simulated, based on a limited test set, and conducted by Anthropic, which has a clear interest in comparing model safeguards. They are evidence about those tests not an independent, universal ranking of model safety.
Still, the broader concern is credible: an open-weight model can be downloaded, adapted, and operated outside a provider’s API controls. That makes safeguards harder to enforce after release, even when the original model includes refusals.
For context on Anthropic’s restricted defensive model strategy, see Project Glasswing and Claude Mythos. The tradeoff is increasingly clear: controlled access can limit misuse, while defenders need timely access to powerful tools to find and fix flaws.
What security teams should do now
The practical response is not panic or a ban on every capable model. It is to assume that exploit development is getting cheaper and adapt the defensive workflow:
- Shorten patch and exposure windows. Prioritize internet-facing and high-impact vulnerabilities, especially when public exploit details or patches appear.
- Test in isolated environments. Keep AI-assisted vulnerability research, exploit reproduction, and patch validation away from production credentials and networks.
- Harden developer and CI systems. Use least privilege, ephemeral credentials, egress controls, and monitoring for unusual tool or process activity.
- Evaluate the whole workflow. Test models, agents, tools, and guardrails against realistic misuse not just direct harmful prompts.
- Use AI defensively with verification. Treat model-generated findings as hypotheses; reproduce and review them before making security decisions.
Our guide to building a Mythos-ready security program covers containment and operational readiness. For the wider acceleration in exploit development, see AI’s shrinking vulnerability-to-exploit timeline.
The takeaway
Anthropic is talking about a competitor, but the strategic signal is bigger than the rivalry: frontier-level cyber capability may be becoming broadly available faster than safety controls can travel with it.
Defenders should focus on resilience rapid patching, strong isolation, tested incident response, and independently verified AI-assisted security work. The model race will continue; the organizations that reduce the time from vulnerability discovery to safe remediation will be better positioned for it.
Frequently Asked Questions (FAQ)
What did Anthropic say about GLM-5.3?
Anthropic reported that GLM-5.3 performed strongly on selected exploit-development evaluations and that its safeguards were bypassed at high rates in the lab’s simulated tests. Anthropic described testing in isolated environments and said the results were specific to its benchmarks and methods.
Can GLM-5.3 create working cyber exploits?
In Anthropic’s reported tests, GLM-5.3 produced end-to-end exploits in 50 of 410 ExploitBench attempts and full control-flow hijacks in 4% of a 100-task internal benchmark. These are benchmark results, not a guarantee of success against arbitrary real-world targets.
How easy is it to bypass GLM-5.3 safeguards?
Anthropic reported engagement rates of 64% under deceptive red-team framing, 92% with prefilling, and 100% after refusal-reduction modification in its simulated tests. These findings are from Anthropic’s evaluation and should not be treated as an independent or universal safety score.
Why is an open-weight cyber model a security concern?
Open weights can be downloaded and run outside the original provider’s API, making it possible for users to modify or remove safeguards. This can increase access for defenders, but it also makes post-release misuse controls more difficult.
What should security teams do about more capable AI exploit models?
Security teams should shorten vulnerability remediation windows, isolate AI-assisted testing, remove production credentials from research environments, restrict network egress, monitor agent actions, and independently verify model-generated findings and patches.
Sources
- Anthropic: GLM-5.3 and the spread of advanced cyber capabilities
- NIST CAISI: Assessment of Z.ai’s GLM-5.3 cyber capabilities
Anthropic’s results summarized here are claims from its own evaluation. Benchmark outcomes depend on test design and should not be generalized to all targets or real-world conditions.