benchmarks
Coding agent clears all 183 ARC-AGI-3 levels
Promtime
benchmarksNVIDIA's AVO, a general purpose coding agent, reached a 100.00 RHAE score on ARC-AGI-3, completing all 183 levels across the benchmark's 25 environments. The figures come from NVIDIA's developer blog and were picked up on Threads. ARC-AGI-3 places agents in turn-based games with hidden rules and no instructions.
At a glance
- AVO runs a continuous loop of inspection, planning, implementation and evaluation, drawing on memory, tools and execution feedback so that each attempt builds on what the previous one produced.
- Beyond the perfect score, AVO needed 12% fewer environment actions than VISTA, while Claude Opus 5 sits at roughly 30% baseline performance on the same interactive benchmark.
- OpenAI has also reported gains on ARC-AGI-3, in its case after enabling reasoning carry-over and compaction in the Responses API, which puts several agent stacks on one scoreboard.
A benchmark that one system clears completely stops discriminating between the systems below it. ARC-AGI-3 was positioned as the harder, interactive round of the ARC series, and a full clear by a general purpose coding agent suggests the format has less headroom than a 30% baseline implies. What likely matters more now is the action count, since efficiency remains measurable after completion does not.
ARC-AGI-3 is an interactive reasoning benchmark: an agent is dropped into a turn-based game, told nothing about its rules, and has to work out the mechanics by acting and observing what changes. The full set spans 25 environments and 183 levels, and AVO's score covers every one of them.
The framing NVIDIA puts on the run goes beyond the score: the developer blog presents AVO as a frontier-level general purpose architecture for long-horizon autonomous agents, that is, a single system meant to hold a task across many steps. The efficiency claim sits alongside it, at 12% fewer environment actions than VISTA over the same benchmark.
The baseline NVIDIA cites for context is Claude Opus 5, at roughly 30% on ARC-AGI-3. OpenAI has separately reported improvements on the benchmark after enabling reasoning carry-over and compaction in the Responses API, which keeps reasoning state across calls rather than restarting each turn.
After a saturated benchmark
NVIDIA has not attached a release or availability date to AVO in the published material, and the post is framed around the benchmark run itself. What remains open is whether the same loop holds up on tasks with longer horizons than a 183 level game set, and whether the ARC Prize team treats a 100.00 as the end of ARC-AGI-3's useful life.
Comments
No comments yet. Be the first.
Join the conversation
Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.
We only use your name and avatar from Google. We never store your email address.
