GLM 5.3 signed a commit as Claude Fable 5

The author of an essay on Future-seems-so-good was running GLM 5.3 in their own harness when it signed a commit "Co-Authored-By: Claude Fable 5". Nothing in the system prompt mentioned Claude or commit conventions. That line is Claude Code's default trailer, and the essay uses the slip to argue that a model's best edit tool is mostly decided in post-training, inside the harness its lab used for RL.
At a glance
- Stencil's February hashline format, which tags every line with a short hash, looked like a broad win, but its headline gain was measured against patch, a format only patch-trained models write reliably.
- The same model can more than double its score by switching harness: Claude Opus 4.5 scored 33.33% on CORE-Bench Hard in HAL's generalist agent and 77.78% through Claude Code.
- The author concedes that most edit-format numbers come from a single benchmark run by hashline's own builders, and that some of the per-model advice in the essay is inference.
If you missed it: in February, Stencil replaced the edit tool with hashline, which tags every line a model reads with a short content hash so that edits point at those tags instead of repeating old text. Stencil's post closed by calling the model the moat and the harness the bridge. After seven months of replications, the author calls that line half wrong, because the bridge is part of the moat.
Against str_replace, hashline adds about 3.4 points on average
The famous jump, Grok Code Fast 1 going from 6.7% to 68.3%, compares hashline with patch. In Stencil's run, Grok 4 failed 50.7% of its patches and GLM-4.7 failed 46.2%. Against Claude-style str_replace, the gain across the 16 models averages about 3.4 points, with a median around 3.5.
The spread runs from +11.3 for Claude Haiku 4.5 and Gemini 3 Flash to -8.3 for DeepSeek V3.2. GPT-5.2 Codex lands at -0.4 and spends 26% more output tokens. Replications split: Trueline cut Claude Code's output tokens by 44%, while a June benchmark of three tools across five models concluded that hashline is almost never the cheapest edit tool.
The telling details are small. Switching tags from 5:af to 5#ZY helped because small models get confused by digits on both sides of the separator. One developer found that a separate hashline_edit tool next to edit gave Haiku-class models 0% success, while the same logic behind a single edit tool gave 95%.
Opus 4.5 scores 77.78% on CORE-Bench Hard inside Claude Code, 33.33% outside
On HAL's CORE-Bench Hard leaderboard, Claude Opus 4.5 scores 33.33% in HAL's generalist agent and 42.22% in CORE-Agent. Nicholas Carlini's scaffold ran it through Claude Code and got 77.78%, then 95.5% after fixing a few grading errors by hand. One row down, the older Opus 4.1 does worse in Claude Code, at 42.22%, than in CORE-Agent, at 51.11%.
Cursor hit the same wall with OpenAI. The Codex model, Cursor wrote, receives a limited set of tools in training and learns to search, read and edit through the shell, so Cursor renamed its tools after shell equivalents like rg. Removing reasoning traces between tool calls dropped GPT-5-Codex by 30% on Cursor Bench, against 3% for mainline GPT-5 on SWE-bench in OpenAI's tests.
OpenAI itself recommends GPT-5-Codex only for agentic coding in Codex or Codex-like environments. When someone post-trained Qwen3-Coder on a real debugger, the tool alone did little: the base model solved 67% of 27 held-out bug tasks without RL and 89% after RL, with median turns falling from 46 to 19.
Claude Code's forgiving client leaves newer Claude models sloppy in other schemas
A tool call is just text. The server flattens the transcript, system prompt and tool definitions into one prompt with marker tokens. The model emits a span the API reads as a call because it was rewarded on that exact format. With constrained decoding, the sampler masks any token that breaks the schema; without it, as in most third-party harnesses, the model follows a learned habit.
Armin Ronacher caught Opus 4.8 and Sonnet 5 inventing keys such as requireUnique, matchCase and oldText2 inside the nested edits[] array of Pi's edit tool, right after closing a long escaped string. In one reproduced session Opus 4.8 failed about 20% of the time. Stripping thinking blocks halved that, and strict tool invocation eliminated it. Codex models showed no such regression.
Ronacher's explanation: Claude Code's client retries when <invoke markup leaks into visible text, repairs broken escapes, accepts old_str and path as aliases and silently drops unknown keys. If RL runs there, a slightly malformed call still earns reward. Picture a driving instructor who never marks you down for skipping the indicator: you pass anyway, and the habit stays with you.
Anthropic counted over 16 million distillation exchanges by February
The author sees two candidate sources for the GLM trailer: distillation from Claude Code trajectories, or pretraining on public GitHub, which now holds many commits Claude Code wrote. One commit can't tell them apart. Because GLM named Fable 5 rather than a generic model, the habit came from recent data. Separately, one HN analysis found K3 calling itself Claude in 7 of 48 samples.
On February 23, Anthropic reported that DeepSeek, Moonshot and MiniMax generated over 16 million exchanges with Claude through about 24,000 fraudulent accounts. MiniMax ran more than 13 million, aimed at agentic coding and tool orchestration, and Moonshot ran 3.4 million. Anthropic's September report, as TechCrunch summarized it, counted nearly 200 million exchanges across five campaigns, the largest, attributed to Alibaba, at 151 million.
The author admits the limits: most edit-format numbers come from one benchmark run by hashline's builders, the public data is small and mostly TypeScript and JavaScript, and neither Anthropic nor OpenAI says outright that it runs RL inside its product harness. In our view, the advice is strongest in its negatives, such as never giving GLM or Grok patch, since the author marks the Claude-shaped recommendation for GLM, Kimi and MiniMax as inference.
Whether labs train on rival harnesses
The tax shrinks if labs train on many harnesses on purpose. Research points that way: SWE-Edit, which treats format choice as a policy to learn, improved edit success by 12.5 points with GRPO training on Qwen3-8B. The author says the incentives currently point the other way. Until that changes, the essay's advice is to measure per model and per language before moving away from a model's native format.
Related stories
- Coding harnesses drove a 70-fold token gap
- FutureOS kept 147 of 178 answers, Codex kept 68
- Jev runs 1,000 code reviews for $0.04, scoring 98%
- Perplexity's agents built a database, then got locked out
- Sai claims a top OSWorld 2.0 score at lower cost
- Microsoft Copilot gets a Code mode on GitHub Copilot tech
Comments
No comments yet. Be the first.
Join the conversation
Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.
