Skip to content

open-models

Xiaomi's new MiMo models carry an unverified top-6 claim

Promtime

Xiaomi put a price tag on its own training run: about 2.62 million dollars for MiMo-V2.6-Pro and about 850,000 dollars for MiMo-V2.6-Flash, each covering 30 reinforcement-learning steps across roughly 750,000 trajectories in fewer than six days. Both models are out with open weights, and the release post on Threads puts Pro at number 6 on the Artificial Analysis Intelligence Index, the best open-weight model on that board.

At a glance

  • Both models are natively omnimodal, aimed at coding, visual tasks and computer use, and live now in AI Studio, MiMo Code, MiMo Desktop, Xiaomi's API platform and OpenRouter.
  • Pricing stays where V2.5 sat: Flash at 0.14 dollars per million uncached input tokens and 0.28 for output, Pro at 0.435 and 0.87, with UltraSpeed ten times that.
  • The rank-6 claim traces to a post dated September 13, 2025, and the claimed parity with Claude Opus 5 and GPT-5.6 Sol we could not verify at all.

If you have not followed MiMo: it began in April 2025 with a 7B model, extending Xiaomi from phones, cars and connected devices into foundation models, and the team is led by Luo Fuli, formerly of DeepSeek and Alibaba's DAMO Academy. According to Wikipedia, the previous MiMo-V2-Pro arrived on 18 March 2026 with over a trillion total parameters, 42 billion active and a one-million-token context window.

Pro scores 46.32 on the Artificial Analysis Intelligence Index v4.3

That score, as reported by Testing Catalog, puts MiMo-V2.6-Pro ahead of Kimi K3 and Qwen3.8 Max and makes it the highest-scoring open-source model in that comparison, by Xiaomi's own reading. The release post also claims parity with Claude Opus 5 and GPT-5.6 Sol across most agent benchmarks.

Pro is the capable one; Flash targets a cheaper balance of intelligence and efficiency. A third variant, Pro-UltraSpeed, promises up to twenty times faster output at the same quality for latency-sensitive work. Xiaomi's own site and OpenRouter confirm that both models shipped.

API prices stay at V2.5 levels

Flash costs 0.14 dollars per million uncached input tokens and 0.28 per million output tokens; Pro is 0.435 and 0.87 on the same basis, according to Testing Catalog. UltraSpeed costs ten times the Pro rates.

MiMo Desktop leaves early access with both models included. Alongside the weights, Xiaomi is publishing the technical report, the training environments and the reinforcement-learning code, framing the launch as a reproducible test of scaled RL and model self-improvement.

Why grade successful runs against each other?

Because a pass/fail check cannot tell a tidy, reliable solution from a wasteful one that stumbled into the right answer. According to RuntimeWire, the MiMo team built graders that compare successful trajectories with one another, rank them and push the reward toward the better paths. Picture a teacher who stops marking homework "done" and starts ranking the ones that are done.

The run used 1,568 prompts per update with 16 rollouts each, producing 3.5 billion to 3.7 billion training tokens per step, and mixed coding, general-agent, visual and cybersecurity tasks into one run rather than a model per category, RuntimeWire writes. Context lengths reached one million tokens. Xiaomi froze the router to limit drift and used adversarial evaluation, anomaly detection and verifier cross-checks against reward hacking.

Vibe World covers Blender assets, a robot arm and a Lean 4 proof

From an image, a video or a text prompt, MiMo-V2.6 can coordinate agents to build interactive 3D scenes and visually test them, create Blender assets, drive a Franka Panda robotic arm from camera feeds, produce frontends and presentations, assemble videos and compose music as scores and MIDI. Xiaomi calls this wider target Vibe World.

Research demonstrations went further: screening materials for capturing PFAS chemicals, and helping formalize a Lean 4 theorem in more than 6,000 lines of kernel-verified code. On the held-out DeepSWE v1.1 test, Flash rose from 48.8 to 65.68 and Pro moved from 58.4 to 72.57. RuntimeWire adds that average pass rates on training tasks improved 25% in relative terms for Flash and 12% for Pro.

Two claims are worth holding at arm's length. The rank-6 line traces to a post dated September 13, 2025, so it reads as a snapshot from over a year ago rather than the current board, and RuntimeWire notes the DeepSWE gains come from Xiaomi's own evaluation, not yet independently reproduced. Oddly, the claim this release leans on hardest is the one with the thinnest paper trail.

What an outside rerun would settle The weights, the technical report, the training environments and the RL code are out together, so the next useful data point is someone outside Xiaomi rerunning DeepSWE v1.1 and landing near 65.68 for Flash and 72.57 for Pro. No date is attached to that, and none is promised. Until then, the cheapest check is Pro on OpenRouter against whatever agent harness you already trust.

Related stories

  1. GLM-5.3 weights land with a hyperscaler clause
  2. Open weights land as a dry run for Qwen4
  3. Cohere claims a WMT26 lead with a non-reasoning translator
  4. A 125B MoE model previews Qwen4's architecture
  5. A 4-bit Mac build fits Qwen3.8-27B in 16.1GB
  6. Zhipu AI will release GLM-5.3 weights in stages after safety reviews

Comments

No comments yet. Be the first.

Join the conversation

Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.

We only use your name and avatar from Google. We never store your email address.