
What happened
Xiaomi's MiMo-V2.6-Pro scored 46 points on Artificial Analysis's Intelligence Index, ahead of Kimi K3 and Qwen, and charges $0.435 per million input tokens and $0.87 per million output tokens — about $0.13 for one test task.
Why it matters
At that price, the model sits on the Pareto frontier of intelligence and cost, meaning buyers can get the strongest openly available model for a fraction of what similarly capable models charge.
What to watch
Xiaomi's claim to the top spot rests on its own reported scores and on Anthropic's unresolved accusation of illegal distillation. Watch whether the roughly 7,000 released training tasks let others reproduce the result.
WHO IT HITSEnterprise AI teams choosing self-hosted or open-weight models can now get frontier-level scores at roughly $0.13 per test task, which changes build-versus-buy math. Legal and compliance teams evaluating those same models, however, now have Anthropic's distillation accusation to weigh.
Summaries like this, in your inbox every morning.
Xiaomi's two-model release pairs a flagship, MiMo-V2.6-Pro, with a smaller, more efficient Flash variant. Pro is a mixture-of-experts model with 1.02 trillion parameters, of which only 42 billion are active per request. Xiaomi says the performance jump came from heavily expanded reinforcement learning — training through trial, feedback, and reward — scaled along three axes: more data per training step, more varied task environments, and more compute for grading solutions. The company says the run took less than six days, and on the DeepSWE coding test Pro climbed from 58.4 to 72.6 while Flash rose from 48.8 to 65.7.
To keep training stable at that scale, Xiaomi froze the model's internal distribution mechanism and added protections against "reward hacking," where a model games its reward rather than solving the task. The released toolkit is unusually open: a technical report, the full training framework, a smaller model for further training, and roughly 7,000 ready-made tasks with automatic graders across software development, cybersecurity, office work, web design, and about 1,000 music composition tasks. Some code comes from real GitHub pull requests by employees and user queries, other task descriptions are model-generated, cyber tasks draw on OSS-Fuzz's collection of real vulnerabilities, and office environments are rebuilt synthetically.
That openness sits against Anthropic's threat intelligence report from two weeks earlier, which examined Claude abuse between December 2025 and August 2026 and named seven Chinese labs. Anthropic says the labs generated about 190 million exchanges to siphon off Claude's capabilities, a practice it calls illegal distillation. In Xiaomi's case, tagged GTG-16008, Anthropic tracked more than 400,000 exchanges over 20 days in March and April 2026, describing Xiaomi as passing user conversations and coding sessions from its own MiMo models through OpenClaw and OpenCode to Claude. The report offers almost no documentation of where the earlier teacher-model training data came from. Whether Xiaomi's top ranking holds up may depend less on the benchmark than on how this dispute is resolved.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, LINE, or Slack.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Teradata is expanding its Tera AI assistant with a Context Engine, a Harness execution system and Agent Skills…
Baselayer raised $35 million in Series A funding led by M13, with Torch Capital, Picus Ventures and Afore Capi…
Heidi Health announced $340 million in new funding — a $100 million Series C led by Blackbird, with Phoenix Co…
Kantata Inc. unveiled Agent Studio, which sits inside its Expertise Engine
Meta's AI agent Muse has leapt to the top of app charts roughly two weeks after its release, helping push up c…

Meta's Muse agent, launched Sept
