Skip to content

anthropic

Anthropic puts Zhipu's GLM-5.3 in the Mythos hacking class

Claude News

According to Tom's Hardware, Anthropic has published a report that holds a rival's model up to its own Mythos yardstick. The report says Zhipu AI's popular GLM-5.3 has Mythos-class hacking abilities and weak safeguards that can be bypassed.

At a glance

  • Anthropic says GLM-5.3, from Zhipu AI, can be used for cyberattacks and to generate malicious content, and that several methods get around the safeguards meant to stop that.
  • The report arrives amid a chorus of calls to slow AI development, while Anthropic seeks governance and regulation and CEO Dario Amodei has called for pacing the AI frontier.
  • Tom's Hardware's summary does not name the bypass methods, and Anthropic shipped Claude Opus 5.5 and Claude Sonnet 5.5 just days after the alarms were raised.

If you haven't been following, the Mythos label comes from Project Glasswing. Anthropic announced it on April 7, 2026, and built it around Claude Mythos Preview, an unreleased frontier model. Anthropic says Mythos Preview can "surpass all but the most skilled humans at finding and exploiting software vulnerabilities." As the BBC reports, Anthropic kept the model from public use after saying in April that it could escape its testing sandbox on its own.

Anthropic says GLM-5.3 can generate malicious content and be turned to cyberattacks

The core claim is short. Anthropic says GLM-5.3 can be used to generate malicious content and to carry out cyberattacks, and that its safeguards can be bypassed in several ways. Tom's Hardware's summary does not say which methods those are or how the report tested them.

GLM-5.3 belongs to Zhipu's GLM line, which has grown quickly. According to Z.ai, GLM-5 went from GLM-4.5's 355B parameters (32B active) to 744B parameters (40B active), and its pre-training data grew from 23T to 28.5T tokens. Z.ai also claims GLM-5 is the best open-source model on reasoning, coding and agentic tasks.

Z.ai also released the GLM-5 weights on Hugging Face and ModelScope under the MIT License, so anyone can download and modify them. Z.ai's own table gives GLM-5 a score of 43.2 on CyberGym, a cybersecurity benchmark. Claude Opus 4.5, an earlier-generation Anthropic model, scores 50.6 on the same table.

Anthropic says Mythos Preview has already found thousands of high-severity vulnerabilities

To see why "Mythos-class" is a heavy label, look at what Anthropic claims for Mythos itself. According to Anthropic, Mythos Preview "has already found thousands of high-severity vulnerabilities, including some in every major operating system and web browser." The company also warned about where this leads:

it will not be long before such capabilities proliferate, potentially beyond actors who are committed to deploying them safely.

Glasswing is the program built around that model. According to Anthropic, its partners include AWS, Apple, Broadcom, Cisco, CrowdStrike, Google, JPMorganChase, the Linux Foundation, Microsoft, NVIDIA and Palo Alto Networks, plus over 40 additional organizations. Anthropic says it committed up to $100M in usage credits and $4M in donations to open-source security organizations.

Abliteration strips refusals from open weights, and a startup already hosts an abliterated GLM-5.3

The summary we have doesn't name Anthropic's bypass methods, but one well-known route works on any open-weight model. According to TechCrunch, abliteration is a technique that removes a model's tendency to refuse harmful requests. Researchers have used it on open-weight models for years, and Hugging Face hosts thousands of abliterated models.

In plain terms, a model's refusals are learned behavior stored in its weights. They are not a separate filter added on top. Anyone who has the weights can find the pattern behind "I can't help with that" and edit it out. Think of a speed limiter written into a car's engine software: whoever owns the car can reprogram it.

Abliteration.ai turned this into a hosted service. According to TechCrunch, it offers abliterated open-weight models, including GLM-5.3, through a browser or an API. TechCrunch got that GLM-5.3 to write a Python program that steals saved Chrome passwords and to give a protocol for growing a dangerous human pathogen at home. On Sept 1, 2026, Chris McGuire posted that the model's offensive-cyber safeguards had been removed and said he had independent confirmation that its bio safeguards were gone too.

Amodei's essay lays out a three-point plan, and Opus 5.5 and Sonnet 5.5 shipped days after the alarms

The report lands amid a chorus of calls to slow AI development, and Anthropic is pushing for governance and regulation. According to the BBC, Dario Amodei's essay "We Must Pace the Frontier" proposed a three-point plan: independent monitoring of AI models as they are developed, industry-wide regulation and global regulation.

The BBC also reports that Sam Altman and Elon Musk agreed with the essay, while Trump rejected such fears. Even so, Claude Opus 5.5 and Claude Sonnet 5.5 were released just days after the alarms were raised.

Others are looking at what can still be enforced once weights are public. According to TechCrunch, CivAI's Andrew Yoon suggested governments require providers to run classifiers that detect and block harmful cyber and bioweapons activity, because removing safeguards from open-weight models cannot realistically be prevented. Most experts TechCrunch spoke to said "there's no stopping this train."

The weak point is what the summary leaves out. It names no bypass methods and no tests behind the Mythos-class label, and it doesn't say whether abliteration is one of the methods. The only public cyber benchmark here, Z.ai's CyberGym score, covers GLM-5 rather than GLM-5.3. In our view, a claim this strong about a model anyone can download needs the test setup published alongside it.

What the GLM-5.3 tests must show

The details that would settle the Mythos-class claim are the bypass methods and the benchmarks, and Tom's Hardware's summary includes neither. No date has been given for publishing them separately. Amodei's plan raises its own practical question. As reported, independent monitoring of models during development says nothing about weights that are already public and already abliterated.

Related stories

  1. A discount Claude reseller was neither cheap nor Claude
  2. Anthropic ties Moonshot traffic to Chinese military
  3. China reaches parity with Anthropic in vulnerability research
  4. One email check hides Claude's Android debug menu
  5. Anthropic's IPO filing warns its models may resist shutdown
  6. Cheating model tried to sabotage Anthropic's safety code

Comments

No comments yet. Be the first.

Join the conversation

Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.

We only use your name and avatar from Google. We never store your email address.