Skip to content

anthropic

Semgrep: GLM-5.2 outperforms Claude Code in IDOR detection

Claude News

The Semgrep security team ran models through their IDOR (Insecure Direct Object Reference) vulnerability benchmark. Zhipu AI's GLM-5.2 achieved a 39% F1 score at a cost of approximately $0.17 per detected vulnerability, surpassing Claude Code running on Opus 4.6 (37%) and Opus 4.8 (28%).

Both runs used identical prompts, but the open-weights GLM-5.2 operated without additional scaffolding, receiving only the code and a hint without endpoint discovery. The Semgrep multimodal pipeline using GPT-5.5 and Opus 4.8 remained ahead at 61% and 53% respectively, specifically due to that scaffolding.

GLM-5.2 is a 750 billion parameter MoE model with approximately 40 billion active parameters per token and a context window of up to 1 million tokens.

Related stories

  1. Anthropic's models miss the frontier in a security PR-review test
  2. GLM-5.2 (max) matches Claude Opus 4.8 on Harvey LAB-AA
  3. China reaches parity with Anthropic in vulnerability research
  4. GLM-5.2 vs. Claude Opus 4.8
  5. Fable 5.1 refuses the knife but heats a gas can anyway
  6. Opus 5 falls to prompt injection 2% of the time

Comments

No comments yet. Be the first.

Join the conversation

Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.

We only use your name and avatar from Google. We never store your email address.