Xiaomi MiMo V2.5 Pro

Xiaomi MiMo V2.5 Pro is Xiaomi’s flagship model with 1.02 trillion total and 42 billion active parameters, designed for frontier reasoning and agentic workflows. Its hybrid attention architecture significantly reduces KV-cache memory, and the context window spans one million tokens. Natively omnimodal for text, image, video, and audio, and fully commercially usable under the MIT license.

Xiaomi Version V2.5-Pro Commercial use permitted MoE 1020 B (42 B active) 1024 K Context 05/2025 $0.435 / $0.87 per 1M

  • Open Weights
  • Frontier
  • OpenRouter
  • Text
  • Vision
  • Video
  • Audio
  • Instruction-Tuned
  • Agentic Orchestrator
  • Batch

Sovereign Risk: MEDIUM Xiaomi is a Chinese company and subject to China’s Data Security Law (DSL) and National Intelligence Law (NIL). The weights are publicly available under the MIT license. When using cloud services, state access to transmitted data is theoretically possible. Local deployment with the public weights reduces the risk — the NIL is only directly relevant when using a cloud API.

Tool-use profile: 6 assets in detail

Comparison of asset performance (P1/P2/Combined) against the fleet average

Asset performance (radar)

Score Breakdown vs. Fleet Average


Tool-use details

Asset performance, reliability, and runtime profile

CrucibleMark evaluates tool use across 6 independent tests. Click a test name for details.

Reliability

  • Tool Call Valid: No
  • Retry: Not required
  • Hallucination: Detected

Reliability measures how consistently a model actually executes tool calls: Tool Call Valid schema and format accepted, Retry Required successful only after retry, Hallucination Flag fabricated tools or parameters detected. All three green means production-ready.

Runtime profile

Call 1
65.39
First Request
MCP
1.19
Protocol Latency
Synthesis
79.6
Response Generation
Total
877.06
Sum of All Phases
Token
17656
Input + Output
Cost
$0
Cost per Run

The runtime profile shows the latency and cost metrics for the model run: Call 1 First Request, MCP Protocol Latency, Synthesis Response Generation, Total sum of all phases, supplemented by Token input and output and Cost cost per run.

Tool-use review

· Instruction-Tuned · Agentic Orchestrator

Deployment Verdict

Conditional deploy, because tool execution is strong, but the model fails to maintain the central trust signal for production MCP pipelines when hallucination is detected and an invalid tool call occurs.

Tool Execution Profile

Xiaomi MiMo V2.5 Pro demonstrates genuine tool intelligence, not just rigid call patterns. In the Web Search & Tool Selection test — which checks the choice between search and direct retrieval without an explicit hint — it correctly identifies the need for web_search. This speaks to robust planning logic in dynamic agent runs. HTTP Fetch & Extract is solid as well.

Performance weakens on the URL Construction test, which measures the autonomous derivation of a target URL and subsequent retrieval. Here the performance is usable, but not deterministic enough for pipelines that derive operational URLs directly from model knowledge. More critical is the finding that the tool call was not consistently valid overall. Since no retry was required, this looks less like a pure formatting issue and more like a situational execution weakness in the orchestration path.

Synthesis Fidelity

How well does it consolidate tool results? Only to a limited degree. The P2 score of 44.17 shows: it can merge retrieved content, but not with consistently sufficient precision. This is especially visible in EU License Research and URL Construction & Fetch, where execution is still usable but consolidation drops off sharply. For production tool pipelines this is a problem, because the actual value only emerges from correctly translating tool data back into reliable statements.

Does it stay within the tool result, or does it fall back on training data? No, not reliably. In the honeypot EU License Research test — which checks whether current license restrictions genuinely come from web sources rather than training knowledge — the model hallucinates. This is not a quality deficiency; it is a security risk. Once a model outputs fabricated or pre-learned facts as a tool result, the entire infrastructure loses its auditability.

Error Resilience

In the 404 test, which checks for transparent behavior on failed retrieval, Xiaomi MiMo V2.5 Pro does not fabricate page content. That is the correct production reflex. However, the communication of the error remains only moderately well consolidated. For operations this is acceptable, because transparency matters more than linguistic quality.

Operational Profile

Call 1: 65.39s. Call 2: 79.60s. MCP latency: 1.19s. Total: 877.06s. Slow. Cost/run: local. Inexpensive to operate, but the runtime is only partially proportionate to the synthesis reliability achieved.

Conclusion & Recommendation

Suitable for agentic research and orchestration pipelines where tool selection, search strategy, and error transparency matter more than the last mile of fact-strict synthesis. Not suitable for compliance, licensing, policy, or other high-trust pipelines where the model must adhere strictly to tool results. If you deploy it, do so with hard response validation, output scoring, and a downstream verification component before any externally consequential decision.

This analysis was generated automatically based on benchmark data. Model used: GPT-5.4 by OpenAI. The raw data and full methodology are documented in the GitHub project.