MiniMax M3

MiniMax M3 is a multimodal MoE model with a context window of one million tokens, focused on agentic workflows, coding, and tool use. Of 428 billion total parameters, only 23 billion are active per token; the model processes text, image, and video as input. Its Chinese origin requires a separate data privacy risk assessment when used via cloud.

MiniMax Version m3 Commercial use permitted MoE 428 B (23 B active) 1000 K Context 05/2026 $0.3 / $1.2 per 1M

  • Open Weights
  • Server
  • OpenRouter
  • Text
  • Vision
  • Video
  • Interactive

Sovereign Risk: HIGH MiniMax is a Chinese company and subject to China’s National Security Law (NSL), which may enable state access to data. The model has been released as open weights, but remains high-risk from a sovereignty perspective when data or workflows are processed under Chinese jurisdiction.

Tool-use profile: 6 assets in detail

Comparison of asset performance (P1/P2/Combined) against the fleet average

Asset performance (radar)

Score Breakdown vs. Fleet Average


Tool-use details

Asset performance, reliability, and runtime profile

CrucibleMark evaluates tool use across 6 independent tests. Click a test name for details.

Reliability

  • Tool Call Valid: No
  • Retry: Not required
  • Hallucination: Detected

Reliability measures how consistently a model actually executes tool calls: Tool Call Valid schema and format accepted, Retry Required successful only after retry, Hallucination Flag fabricated tools or parameters detected. All three green means production-ready.

Runtime profile

Call 1
1.84
First Request
MCP
0.91
Protocol Latency
Synthesis
10.91
Response Generation
Total
81.99
Sum of All Phases
Token
16951
Input + Output
Cost
$0
Cost per Run

The runtime profile shows the latency and cost metrics for the model run: Call 1 First Request, MCP Protocol Latency, Synthesis Response Generation, Total sum of all phases, supplemented by Token input and output and Cost cost per run.

Tool-use review

Updated

Deployment Verdict

Conditional deploy, because tool execution is strong, but one invalid tool call and a detected hallucination limit confidence in fully autonomous MCP pipelines.

Tool Execution Profile

MiniMax M3 demonstrates genuine tool intelligence rather than mere pattern-following. On the Web Search and Tool Selection test — which checks whether the model distinguishes between search and direct retrieval without being prompted — it selects the correct tool confidently. This points to workable planning logic in agentic workflows. It also executes the necessary steps reliably on EU License Research and Multilingual Search & Synthesis.

The weakness lies not in tool selection but in execution consistency. On the URL Construction test, which measures correct URL derivation followed by a fetch, it achieves only limited precision in the target address and then loses ground significantly in result utilization. There is also a hard protocol signal: at least one tool call was not valid. For MCP operation, this means the planning side is sound, but the last mile requires guardrails. Retry was not needed, so this is not a mere formatting issue under load — it is a genuine reliability finding.

Synthesis Fidelity

How well does it condense tool results? Only partially reliable. P2 performance is the clear weak point of this run. Solid on EU License Research, HTTP Fetch & Extract, and Web Search & Tool Selection, but a marked drop on the URL Construction test shows that correctly triggered retrievals do not consistently produce precise, decision-ready summaries. This matters for production pipelines because it is the condensed output — not the retrieval itself — that feeds downstream decisions.

Does it stay within the tool result or fall back on training data? On the Honeypot EU License Research task, it stays within the tool path and does not answer from prior world knowledge. That is the more important trust signal. At the same time, the global hallucination finding remains a safety risk: once a model presents fabricated facts as a tool result, it undermines the reliability of the entire infrastructure.

Error Resilience

Acceptable for production. On the 404 test — which checks whether a failed tool call is openly reported or papered over with invented page content — MiniMax M3 communicates the failure transparently. It does not hallucinate substitute content. This is the minimum requirement for robust tool pipelines, and the model meets it here.

Sovereignty Profile

Locally deployable and broadly fleet-competitive, but not sovereignty-leading. The Sovereignty Gap sits at -0.89 points below the fleet average of 68.17. In practice: local deployability is a genuine advantage, and the performance gap relative to the broader fleet is narrow. Jurisdictional risk in cloud use remains high due to provenance and must be assessed separately.

Conclusion & Recommendation

Suitable for agentic research and orchestration pipelines with upstream tool validation, schema checking, and a second instance for result sign-off. Not suitable for fully autonomous compliance, policy, or fact systems where a single invalid tool call or a fabricated summary feeds directly into decisions. As a local model, it is attractive if you weight tool selection and error transparency more heavily than synthesis-precise final answers.

This analysis was generated automatically based on benchmark data. Model used: GPT-5.4 by OpenAI. The raw data and full methodology are documented in the GitHub project.