model-releases
Diffusion model claims 1,107 tok/s in Mercury 2.5
Promtime
model-releasesInception has launched Mercury 2.5, a diffusion language model the lab puts at 1,107 tokens per second on widely available NVIDIA GPUs. The launch was announced on September 8, 2026, with the performance figures relayed in a post on Threads.
At a glance
- Diffusion language models such as Mercury generate and refine multiple token positions in parallel, instead of producing text strictly one token at a time the way autoregressive systems do.
- Inception reports two headline numbers for Mercury 2.5: 1,107 tokens per second on widely available NVIDIA GPUs, and a 40% gain in intelligence over its previous model, Mercury 2.
- Inception positions Mercury 2.5 as matching Gemini 3.5 Flash and GPT 5.6 Luna Low at much higher speed, a comparison resting on its own launch post rather than independent benchmarks.
Throughput is where diffusion approaches are supposed to pay off, and Mercury 2.5 reads as an attempt to convert that structural advantage into parity on quality rather than a niche speed record. If the comparisons hold, the interesting part is not the token rate itself but the class of workloads it opens up: latency-bound coding assistants, agent loops and interactive tools where autoregressive decoding sets the ceiling. That verification is still outstanding.
The parity claim covers two models. Inception says Mercury 2.5 performs on par with Gemini 3.5 Flash and with GPT 5.6 Luna Low while generating text considerably faster. Those head-to-head results come from the September 8 launch post and have not been backed by outside benchmarks.
Mercury runs on a diffusion approach: the model generates and refines multiple token positions in parallel rather than producing text strictly one token at a time. Inception puts Mercury 2.5 throughput at 1,107 tokens per second on widely available NVIDIA GPUs.
Most widely used large language models decode autoregressively, emitting one token after the previous one, so generation time scales with output length. Mercury is Inception's line of models built on the diffusion approach instead, with Mercury 2.5 as the current release.
The second number in the launch is a 40% gain in intelligence over Mercury 2, the previous model in the same family. Inception announced Mercury 2.5 on September 8, 2026, and both headline figures were published in the launch post that accompanied it.
Untested against outside benchmarks
Neither figure comes with its conditions spelled out: the source material does not name the benchmark suite behind the 40% intelligence gain, the specific NVIDIA GPUs used for the 1,107 tokens per second measurement, or a price for Mercury 2.5. Independent evaluation of the parity claim against Gemini 3.5 Flash and GPT 5.6 Luna Low remains the open item.
Comments
No comments yet. Be the first.
Join the conversation
Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.
We only use your name and avatar from Google. We never store your email address.
