Xiaomi MiMo V2.6 Flash

The smaller sibling of the MiMo-V2.6 series prioritizes efficiency over maximum size: MiMo-V2.6-Flash-RL by Xiaomi activates around 15 billion of 309 billion parameters per token, matching the flagship Pro on agent and coding tasks at roughly one-third of the API price. Omnimodal for text, image, video, and audio, context up to 1 million tokens, Open Weights under the MIT license.

Xiaomi Version V2.6-Flash Commercial use permitted MoE 309 B (15 B active) 1024 K Context $0.14 / $0.28 per 1M

  • Open Weights
  • Server
  • OpenRouter
  • Text
  • Vision
  • Video
  • Audio
  • Agentic Orchestrator
  • Long Context
  • Interactive

Sovereign Risk: MEDIUM Xiaomi releases MiMo-V2.6-Flash-RL under MIT with fully open weights, which significantly improves operational provenance for local deployment. As a Chinese developer, however, Xiaomi remains subject to national laws (including the National Intelligence Law and the Data Security Law), which remains relevant when using the model via Xiaomi’s own API platform; with purely local self-hosting, the operational risk is substantially reduced.[434][445]

Key metrics

Score · Latency · Cost · Quality

Total Score Silver
75.04
Routine
45.33
Reasoning
29.71

Rank #33

LLM Judge Avg
3.79
100 Coverage
Avg Task Duration
41.55
Interactive
Token Rate
43.24
Output Rate
P95 Latency
121.74
Top 5 %
Total Tokens
117400
Output Volume
Cost per 1K
$0.0003
USD / 1K Requests
Benchmark Cost
$0.03
Total · 117400 tok

Benchmark modules

10 modules · weighted · vs. model median & top performer

Xiaomi MiMo V2.6 Flash Best model Ø All models
Code Quality 78.16
CLI Benchmark 90.67
Logical Reasoning 75.64
UX Writing 69.87
Documentation 79.75
Content Transform. 70
Cultural Intelligence 70.64
Synthesis Quality 59.17
Tool Execution 90
ToolUse Score 73.42
Benchmark Cost $0.03

Token efficiency & latency

Consumption per module vs. model median

Token consumption per module

Performance profile