Skip to content

google

Gemini 4 Argon wins at office work and trails in coding

Promtime

Google's own benchmark table shows Gemini 4 Argon fully completing only about one in five tasks on Harvey's Legal Agent Benchmark. That 19.6% is still nearly triple Claude Fable 5.1's score. As The New Stack reports, the new flagship leads or ties in 13 of 18 tests against OpenAI and Anthropic, yet only a small group of testers can use it.

At a glance

  • Google has announced Gemini 4 Argon, a flagship that effectively replaces the Gemini 3.5 Pro promised at I/O in May, and has opened it only to testers in its Fairwind Program.
  • Argon's widest margins come in knowledge work and long context, including 51.3% on Zapier's AutomationBench, while its coding results run from a record 77.9% on DeepSWE v1.1 to last place on two tests.
  • Pricing starts at $2 per million input and $10 per million output tokens, then rises to $4 and $20, and Google has given no date for access beyond Fairwind.

In case you missed it, Google announced Gemini 3.5 Pro at I/O in May for a June launch, and Flash models came instead. According to Codersera Blogs, 3.5 Pro missed targets in June, mid-July and early August, and 9to5Google says Argon follows its cancellation under a new naming scheme. The launch came a day after Sundar Pichai co-signed a commitment to "self-police" with Anthropic, Meta, Nvidia, OpenAI and SpaceX after meeting President Donald Trump.

Argon leads or ties in 13 of 18 tests, and its legal score is nearly triple Fable 5.1's

Google's table pits Argon against OpenAI's GPT-6 Astra and Anthropic's Claude Opus 5.5 and Claude Fable 5.1 in 18 tests, and Argon takes first place, outright or tied, in 13. The big wins cluster in knowledge work. On Zapier's AutomationBench it scores 51.3% to 42.5% for Opus 5.5, and on Vals Finance Agent v2 it posts 65.4%, with Fable 5.1 next at 58.9%.

On Harvey's Legal Agent Benchmark, Argon's 19.6% compares with 6.7% for Fable 5.1, 5.4% for GPT-6 Astra and 3.8% for Opus 5.5. Long context is the other strong area. On GraphWalks with inputs from 256K to 1M tokens, Argon reaches 84.2%, more than 12 points ahead of GPT-6 Astra at 71.8%.

Elsewhere the leads are thin. Argon tops the Vals Index, Agent's Last Exam, Chartography and GraphWalks up to 128K by under two points each. On the Vals Index, for example, it scores 68.9% to 67.0% for Opus 5.5. On LVBench, a long-video test, it leads with 91.7% against 87.5% for GPT-6 Astra.

Coding splits between a record 77.9% on DeepSWE v1.1 and last place on two other tests

Google calls Argon's 77.9% on DeepSWE v1.1 a new state of the art, ahead of Opus 5.5 at 74.2% and GPT-6 Astra at 74.1%. Argon also leads Vibe Code Bench with 91.9%, though all four models score above 89% there.

The other two coding tests go the opposite way. Argon finishes last on FrontierSWE v2 with 55.0%, 10.5 points behind GPT-6 Astra, and last on Terminal-Bench 4.0 with 57.4%, nine points behind Opus 5.5. It also trails Opus 5.5 on PostTrainBench, 45.3% to 49.3%, GPT-6 Astra on Terminal-Bench Science 0.1, 57.6% to 68.1%, and GPT-6 Astra on the OSWorld-2.0 offline subset, 69.2% to 72.6%.

On CWE-bench v1, the only cyber test in the post with rival scores, Argon ties GPT-6 Astra and xAI's Grok 4.7 at 68%, with Opus 5.5 at 67%. The OpenAI and Anthropic models ran in their own agent harnesses, Codex and Claude Code, so that leaderboard measures each model and its tooling together.

Google says Argon agents freed more than 300 TiB of data-center memory

Google says it trained Argon to find, validate and patch software vulnerabilities autonomously, and it is releasing the model without cyber guardrails to Fairwind participants and its own teams. Wiz, which Google acquired for $32 billion in March, uses Argon in its Scan for Good initiative. According to Google, the model found a critical vulnerability in healthcare software used by hospitals worldwide that earlier frontier models had missed.

Testing Catalog adds that the flaw exposed sensitive personal information, and it relays other internal examples. Thousands of Google employees use Argon for coding, research and writing. The model beat a published quantum algorithm optimization baseline by 40% within minutes, and Argon agents found memory optimizations that freed more than 300 TiB after rollout, with 500 TiB to 1 PiB projected in total.

According to the same report, Argon agents are supporting C and C++ migrations to Rust. The targets range from core libraries with tens of thousands of lines to the Fuchsia Zircon kernel at more than 800K lines, and every migration goes through audits, emulation testing and review before production. In a Rust port of the libgav1 decoder, Argon replaced 32K lines of SIMD code with safe Rust that ran 2.7 times faster.

Argon costs $2 and $10 per million tokens at first, then $4 and $20

During an introductory period, Argon costs $2 per million input tokens and $10 per million output tokens. Testing Catalog reports that cached input costs 95% less than regular input. Afterward the price rises to $4 and $20, and that later output rate matches the $20 per million output tokens Anthropic charges for Opus 5.5.

Testing Catalog also quotes outside rankings. Artificial Analysis says Argon equals GPT-6 Astra on its Intelligence Index at 60% of the cost per task with discounted prices. Arena put Gemini 4 Argon (High) at #1 in Text Arena with 1525 points and #8 in Code Arena: WebDev with 1679, at a blended $8 per million tokens. A Wall St Engine post it cites showed Alphabet shares up 3%.

Argon can generate up to one million output tokens, up from 64,000

One million input tokens is now standard for frontier models, but most competitors are not pushing output limits forward. Previous Gemini models could generate up to 64,000 output tokens, and Argon can generate up to one million. Google explains the change in its announcement:

When the model has the headroom to think deeply and generate hundreds of thousands of tokens in a single trajectory, it adds a new level of depth in reasoning to solve tough problems in one go.

Think of it as a whiteboard. A reasoning model works through a problem by writing out its intermediate steps as tokens, and the output cap decides how much board it gets before it has to stop. A bigger board lets a single trajectory carry hundreds of thousands of tokens of working.

Access also comes in stages. Testing Catalog describes the Fairwind Program as open to trusted cyber defenders. Google is collecting their feedback, iterating on guardrails and taking part in the U.S. government's voluntary process for pre-release model access. Paid API customers and AI Ultra subscribers come next, then developers, enterprises and consumers.

All of these benchmarks come from Google, and the cyber claims have no outside yardstick. Argon's 85.8% on Google's internal vulnerability discovery test and 70.9% on Wiz's penetration testing benchmark are compared only with Gemini 3.8 Flash Cyber, which scored 71.0% and 58.2%. In our view, the million-token cap needs some arithmetic: at the later $20 per million output tokens, one response that fills it costs $20 in output alone.

When Argon leaves Fairwind

Google has not given a date for the next step, when paid API customers and Google AI Ultra subscribers get access. It also has not said how long the $2 and $10 introductory pricing lasts before the move to $4 and $20. Nor has it said whether the version without cyber guardrails will reach anyone outside Fairwind and Google's own teams.

Related stories

  1. Gemini 3.8 Live Extended Thinking tops GPT Live 1
  2. Cohere Embed 5 splits indexing and search across two models
  3. Claude Sonnet 5.5 keeps Sonnet 5's price, cuts cost per task
  4. Gemini 3.8 Flash TTS takes stage directions line by line
  5. Xiaomi's new MiMo models carry an unverified top-6 claim
  6. xAI's new transcriber marks speakers and drops the ums

Comments

No comments yet. Be the first.

Join the conversation

Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.

We only use your name and avatar from Google. We never store your email address.