Xiaomi MiMo V2.5

Xiaomi MiMo V2.5 is a natively omnimodal MoE model with 310 billion total and 15 billion active parameters. The model processes text, image, video, and audio within a single architecture; the context window supports up to one million tokens. Fully commercially usable under the MIT license, from a Chinese manufacturer jurisdiction with a corresponding assessment for cloud deployment.

Xiaomi Version V2.5 Commercial use permitted MoE 310 B (15 B active) 1024 K Context 05/2025 $0.14 / $0.28 per 1M

  • Open Weights
  • Server
  • OpenRouter
  • Text
  • Vision
  • Video
  • Audio
  • Instruction-Tuned
  • Agentic Orchestrator
  • Interactive

Sovereign Risk: MEDIUM Xiaomi is a Chinese company and subject to China’s Data Security Law (DSL) and National Intelligence Law (NIL). The weights are publicly available under the MIT license. When using cloud services, state access to transmitted data is theoretically possible. Purely local deployment with the public weights reduces the risk — the NIL is only directly relevant when using cloud APIs.

Tool-use profile: 6 assets in detail

Comparison of asset performance (P1/P2/Combined) against the fleet average

Asset performance (radar)

Score Breakdown vs. Fleet Average


Tool-use details

Asset performance, reliability, and runtime profile

CrucibleMark evaluates tool use across 6 independent tests. Click a test name for details.

Reliability

  • Tool Call Valid: No
  • Retry: Not required
  • Hallucination: Detected

Reliability measures how consistently a model actually executes tool calls: Tool Call Valid schema and format accepted, Retry Required successful only after retry, Hallucination Flag fabricated tools or parameters detected. All three green means production-ready.

Runtime profile

Call 1
13.41
First Request
MCP
1.03
Protocol Latency
Synthesis
36.92
Response Generation
Total
308.18
Sum of All Phases
Token
12362
Input + Output
Cost
$0
Cost per Run

The runtime profile shows the latency and cost metrics for the model run: Call 1 First Request, MCP Protocol Latency, Synthesis Response Generation, Total sum of all phases, supplemented by Token input and output and Cost cost per run.

Tool-use review

· Instruction-Tuned · Agentic Orchestrator

Deployment Verdict

Conditional deploy, because tool execution is strong, but invalid tool calls and a hallucination signal limit confidence in an unsupervised MCP pipeline. The Combined Score is good, but does not serve as a free pass here.

Tool Execution Profile

Xiaomi MiMo V2.5 demonstrates clear tool intelligence in selecting the correct path. In the Web Search & Tool Selection test — which checks whether the model searches first rather than fetching directly, without any hint — the model correctly identifies the need and achieves full tool execution. This argues against a rigid pattern. In the URL Construction test, which measures the precise derivation of a target URL and the subsequent fetch, it is serviceable but not deterministic enough. A P1 of 80 means in practice: the model reaches the goal frequently, but not reliably enough for pipelines with strict schema adherence.

The finding “Tool call valid: false” is critical. This is not merely a stylistic error but an integration risk at the MCP level. On the positive side, no retry was required. This points more toward isolated protocol or argument errors than a fundamental comprehension problem.

Synthesis Fidelity

How well does it consolidate tool results? Only adequately. A P2 of 70 appears acceptable, but variance is high. HTTP Fetch & Extract and Tool Failure Handling (404) are solid at 80. By contrast, EU License Research at 40, Web Search & Tool Selection at 40, and especially Multilingual Search & Synthesis at 15 fall significantly short. The model can retrieve information, but does not consistently condense it into reliable final answers.

Does it stay within the tool result or fall back on training data? In the honeypot EU License Research test — designed to verify whether current license restrictions genuinely originate from web sources — no hallucination was detected. This is the most important exonerating data point. At the same time, “Hallucination detected: true” appears globally. This leaves a residual security risk: once a model presents fabricated facts as tool results, the entire tool infrastructure becomes questionable.

Error Resilience

In the 404 test — which measures transparent behavior on tool failures rather than fabricated page content — MiMo V2.5 responds in a production-ready manner. A P2 of 80 and no hallucination despite the error mean: it communicates failures openly and does not fill gaps with substitute facts. For productive agents, this is a viable baseline.

Sovereignty Profile

Locally deployable and therefore attractive for sovereign deployments. On the performance side, with a Sovereignty Gap of -0.89 points below the fleet average of 68.17, it sits practically at fleet level.

Conclusion & Recommendation

Suitable for local, sovereign research and orchestration pipelines with human review or downstream validation. Less suitable for compliance, policy, or multilingual synthesis workflows in which the final answer itself must be evidentially sound. If you deploy MiMo V2.5, do so as a tool operator with a tight output schema and an external verifier — not as the final trusted authority.

This analysis was generated automatically based on benchmark data. Model used: GPT-5.4 by OpenAI. The raw data and full methodology are documented in the GitHub project.