> ## Content Index
> Fetch the complete content index at: https://www.metatalks.ai/llms.txt
> Use this file to discover other available public pages before exploring further.

# Open-weight GLM-5.3 nears Mythos Preview at building exploits, Anthropic finds
- URL: https://www.metatalks.ai/open-weight-glm-5-3-nears-mythos-preview-at-building-exploits/
- Published: 2026-09-30T14:44:00.000Z
- Updated: 2026-09-30T14:43:59.000Z
- Author: Al
- Tags: News, AI Security, AI safety, #newswire

**Anthropic's red team got past its safeguards in 64% to 100% of simulated attacks, and copies with its refusals stripped out were public within days.**

GLM-5.3, Zhipu AI's newest open-weight model, built working end-to-end cyber exploits almost as often as Claude Mythos Preview in [simulated tests](https://www.anthropic.com/research/glm-5-3-and-the-spread-of-advanced-cyber-capabilities?ref=metatalks.ai) run by Anthropic's Frontier Red Team. On ExploitBench, which uses known flaws in Chrome's V8 engine, it succeeded in 50 of 410 attempts, against 56 for Mythos Preview.

Mythos Preview was billed in [Anthropic's announcement](https://www.anthropic.com/glasswing?ref=metatalks.ai) five months ago as the first model able to build such exploits on its own. The red team concludes that Zhipu, which goes by Z.ai outside China, released a model close to it without meaningful protection against misuse.

In one hands-on session with a human expert, GLM-5.3 spent a day finding several previously unknown flaws in the JavaScript engine of a widely used browser's Linux build. It chained them into a web page that reads arbitrary files from a visitor's machine. In another, the smaller GLM-5.3-Flash turned a disclosed Chrome flaw and one other known bug into a reliable exploit chain with little guidance. It took 20 minutes of a person's time and 8 hours of model work, and would have cost $20.40 at Zhipu's API prices.

Out of the box, GLM-5.3 refused openly hostile requests to attack critical systems. Telling it that it was a red-team agent in an exercise got it to engage 64% of the time, and writing the start of its reasoning so it seemed to have already agreed raised that to 92%.

Abliteration, a standard technique for cutting refusals, worked every time, and because the weights are open, anyone can apply it. Refusal rates fell from above 90% to between 2% and 12% on three safety benchmarks, while cyber-task scores slipped only a few percent. An experienced team would need about $1,200 of compute, the red team estimates. Several developers had posted abliterated copies within days of the release.

On Claude, Anthropic says, its safeguards stopped the deceptive prompts, and the other two routes are closed: its API offers no way to write the start of Claude's reasoning, and its weights are not released.

On Sept. 17, CAISI, the AI center within NIST, released [its own evaluation](https://www.nist.gov/news-events/news/2026/09/caisis-assessment-zais-glm-53-cyber-capabilities?ref=metatalks.ai). It found no released open-weight model stronger in cyber and put GLM-5.3 about four months behind the leading US models. Anthropic's findings broadly match. But CAISI tested those US models with their cyber safeguards off where relevant, and some are offered only to vetted users. Anyone can download GLM-5.3.

Defenders can use the same skills, Anthropic grants. Hugging Face did so with the previous version: in July, when commercial models' guardrails blocked the attack commands its responders needed to analyze after a breach, it [ran the forensics](https://huggingface.co/blog/security-incident-july-2026?ref=metatalks.ai) on GLM-5.2 on its own servers.

Anthropic wants governments to safety-test sufficiently capable models, including GLM-5.3's successors. It published the analysis the day Trump hosted AI executives, Anthropic CEO Dario Amodei among them, at a [White House lunch](https://www.cnbc.com/2026/09/29/tech-white-house-ai-lunch-trump.html?ref=metatalks.ai). Afterward, Trump said he and the tech leaders had signed a "morally binding" AI document, and he argued for self-regulation by the industry.