Xiaomi MiMo V2.5

Xiaomi MiMo V2.5 is a natively omnimodal MoE model with 310 billion total and 15 billion active parameters. The model processes text, image, video, and audio within a single architecture; the context window supports up to one million tokens. Fully commercially usable under the MIT license, from a Chinese manufacturer jurisdiction with a corresponding assessment for cloud deployment.

Xiaomi Version V2.5 Commercial use permitted MoE 310 B (15 B active) 1024 K Context 05/2025 $0.4 / $2 per 1M

  • Open Weights
  • Frontier
  • OpenRouter
  • Text
  • Vision
  • Video
  • Audio
  • Instruction-Tuned
  • Agentic Orchestrator
  • Interactive

Sovereign Risk: MEDIUM Xiaomi is a Chinese company and subject to China’s Data Security Law (DSL) and National Intelligence Law (NIL). The weights are publicly available under the MIT license. When using cloud services, state access to transmitted data is theoretically possible. Purely local deployment with the public weights reduces the risk — the NIL is only directly relevant when using cloud APIs.

Tool-use profile: 6 assets in detail

Comparison of asset performance (P1/P2/Combined) against the fleet average

Asset performance (radar)

Score Breakdown vs. Fleet Average


Tool-use details

Asset performance, reliability, and runtime profile

CrucibleMark evaluates tool use across 6 independent tests. Click a test name for details.

Reliability

  • Tool Call Valid: No
  • Retry: Not required
  • Hallucination: Not detected

Reliability measures how consistently a model actually executes tool calls: Tool Call Valid schema and format accepted, Retry Required successful only after retry, Hallucination Flag fabricated tools or parameters detected. All three green means production-ready.

Runtime profile

Call 1
4.16
First Request
MCP
1.22
Protocol Latency
Synthesis
12.63
Response Generation
Total
108.08
Sum of All Phases
Token
14719
Input + Output
Cost
$0
Cost per Run

The runtime profile shows the latency and cost metrics for the model run: Call 1 First Request, MCP Protocol Latency, Synthesis Response Generation, Total sum of all phases, supplemented by Token input and output and Cost cost per run.

Tool-use review

Updated · Instruction-Tuned · Agentic Orchestrator

Deployment Verdict

Deploy conditionally, because tool execution is strong, but synthesis does not adhere reliably enough to tool results, and tool calls within the run were not consistently valid.

Tool Execution Profile

Xiaomi MiMo V2.5 shows clear orchestration strength. In the Web Search & Tool Selection test, it recognizes without an explicit hint that a web search is required before a direct fetch — evidence of genuine tool selection rather than a rigid fetch-first pattern. In the URL Construction test, it derives the target URL on its own and executes the retrieval usably, though not with the precision expected for deterministic pipelines. P1 across the full suite is strong. The model plans well, but does not consistently produce protocol-clean tool calls. For MCP infrastructures, this means: usable as an orchestrator, but still in need of safeguards as a strictly formal tool interface.

Synthesis Fidelity

How well does it condense tool results? Only adequately. A P2 score of 70 is not a failure, but it is too inconsistent for a frontier agent model. Solid on HTTP Fetch & Extract and URL Construction & Fetch, noticeably weaker on EU License Research and Multilingual Search & Synthesis. The pattern is clear: it can pull together facts from a single retrieval usably, but when research tasks are ambiguous or cross-jurisdictional, condensation loses precision.

Does it stay within the tool result or fall back on training data? The honeypot for EU License Research — which tests whether current license restrictions are actually drawn from web sources — ends without a detected hallucination. That is the critical trust point. At the same time, P2 there is only 40. The model does not fabricate anything obvious, but it does not hold the retrieved findings together sharply enough. For compliance-adjacent responses, that is too soft.

Error Resilience

In the 404 test, which measures transparent error communication rather than fabricated page content, the model stays on the safe side. It does not hallucinate replacement content despite a failed retrieval. A P2 of 60 indicates, however, that the error communication is not always concise and operationally clean enough. For production use, this is acceptable. The behavior does not break tool integrity.

Operational Profile

Call 1: 4.16s. MCP latency: 1.22s. Call 2: 12.63s. Total: 108.08s. Slow for the quality delivered. Cost/run: local. Economical to operate when the necessary hardware is already in place.

Conclusion & Recommendation

Suitable for agentic pipelines where tool selection, multi-step research, and fault-tolerant orchestration matter more than high-precision final synthesis. Also suitable for local or controlled deployments where open weights and multimodality are priorities. Not the first choice for compliance, policy, licensing, or executive summary pipelines where the response must remain strictly and tightly grounded in tool findings. In those cases, a downstream verifier or a stronger synthesis model should handle the final response.

This analysis was generated automatically based on benchmark data. Model used: GPT 4.5 by OpenAI. The raw data and full methodology are documented in the GitHub project.