MiniMax M3

MiniMax M3 is a multimodal MoE model with a context window of one million tokens, focused on agentic workflows, coding, and tool use. Of 428 billion total parameters, only 23 billion are active per token; the model processes text, image, and video as input. Its Chinese origin requires a separate data privacy risk assessment when used via cloud.

MiniMax Version m3 Commercial use permitted MoE 428 B (23 B active) 1000 K Context 05/2026 $0.3 / $1.2 per 1M

  • Open Weights
  • Server
  • OpenRouter
  • Text
  • Vision
  • Video
  • Interactive

Sovereign Risk: HIGH MiniMax is a Chinese company and subject to China’s National Security Law (NSL), which may enable state access to data. The model has been released as open weights, but remains high-risk from a sovereignty perspective when data or workflows are processed under Chinese jurisdiction.

LLM Model Review

Created on

MiniMax M3 achieves an overall score of 80.19 percent and carries the speed profile Interactive DevOps Expert. This is not a model for casual conversation, but a serious workhorse: fast enough for interactive loops, broad enough for everyday use, capable enough to hold its own in code and tooling contexts without coming across as a copywriter in disguise. At the same time, this run reveals the architecture’s character with considerable clarity: an agentically oriented Frontier model with vision capabilities and optional thinking depth, tested here in the standard mode of the cloud endpoint without the thinking toggle — meaning exactly what a regular API user actually gets. Sovereign Risk: HIGH — MiniMax is a company based in China; cloud usage falls under Chinese legal frameworks, and the reviewed provider documentation lists no GDPR-compliant DPA.

Header Notes: Stability and Reliability

Metric Value Rating Analysis
Timeout Rate 1/49 Sporadic The model shows sporadic dropouts that would require retries in practice. For a cloud Open Weights model via OpenRouter, this is not a device limit but an API and routing risk.
P95 Response Time 70.77 s Problematic Significant outliers that disrupt workflow. In five percent of cases, interaction turns into waiting.

Architecture and Classification

The preliminary classification hits the mark surprisingly well. MiniMax M3 is labeled General, but its actual temperament is agentic. You notice this in how often it not only completes tasks but structures, prioritizes, and enriches them with actionable steps. Add to that Vision-Capable. This benchmark, however, measures almost exclusively text-based work. That means we are only seeing a slice of the platform’s full capability here. Anyone purchasing the model for image or video understanding will get only half the picture from these results.

More important is the second axis: Thinking-Optional. This test run operated in mode n/a, meaning without a visible toggle for extended reasoning. The model was evaluated in its default behavior. That is precisely what makes the result relevant. It does not show the best possible demo configuration, but the out-of-the-box state. For readers, this means: the reasoning performance here is not the product of an explicitly unlocked thinking budget, but what the cloud endpoint delivers by default.

The model class also sets the bar high. MiniMax M3 is a Frontier model, competing in the most expensive and demanding tier of the benchmark. Its MoE architecture with 428 billion total parameters and 23 billion active parameters is not a pure muscle mass, but more of a well-organized team of specialists. It should therefore be evaluated not by the imposing total figure, but by the actual active capacity per token. That is precisely where it delivers a remarkable amount.

Performance Profile: Fast, but Not Always on Time

The speed profile Interactive DevOps Expert is more than a nice label. It describes a user type: models that respond quickly enough in shell, code, diagnostics, and tooling loops to avoid turning every iteration step into a coffee run. MiniMax M3 fulfills this profile credibly. The measured generation speed in cloud operation via OpenRouter is clearly within the interactive range.

That speed needs to be put in proper context, however. For a Cloud Open Weights model, this is never an abstract value of the model alone — it is always also a benchmark of the deployed infrastructure and the provider’s network path. In other words, the pace belongs not only to MiniMax M3, but also to the stack delivering it. Impressive nonetheless.

The flaw lies in the tail. The outliers are real, and they are a nuisance for productive agent chains. Anyone integrating a model into a workflow where multiple tool calls, verification steps, and follow-up queries cascade needs reliability almost more urgently than raw speed. MiniMax M3 is good here, but not flawless. A single timeout is not a disaster. It is simply the kind of small defect that looks large at night inside automations.

Code Quality: A Security Auditor with a Useful Toolkit

In the Code Quality Audit, MiniMax M3 plays its strongest hand. The security analysis of the flawed PHP code is not only complete but professionally prioritized. The model identifies more vulnerabilities than the reference list without descending into mere name-dropping. More importantly, it spots the genuinely nasty issues. Mail header injection via CRLF, type juggling in API authentication, predictable reset tokens, session fixation, and header injection are not just named but correctly explained and accompanied by production-ready fixes.

That is the difference between a model that has learned buzzwords and one that understands security patterns. mysqli_prepare, bind_param, password_hash(), hash_equals(), random_bytes(), session_regenerate_id(true) are not decoration here — they are appropriately applied tools. MiniMax M3 does not write security fan fiction. It repairs.

The agentic orientation helps here. The model sorts by criticality, adds layers of detail, and supplements with immediate action items. In editorial terms: it delivers not just a diagnosis but a usable set of instructions. That is exactly what you want from a Frontier model with a tool and workflow focus.

The minor blemish is verbosity. In Code Quality, MiniMax M3 produces on average significantly more text than the fleet median. This does not diminish quality, but it does increase API costs for the same substantive content. For individual queries, that is irrelevant. At high request volumes, it adds up quickly.

CLI and Tool Use: Strong Hand, but Not Blindly Trustworthy

The CLI domain is clearly among the model’s strengths. The benchmark attests a very high level here, and that fits MiniMax M3’s character. It is fast on operational tasks, structures workflows sensibly, and comes across in DevOps-adjacent scenarios like an assistant who knows the workspace and does not need to look for the inbox first.

This is also, however, where the most serious flaw of the entire run sits. In one tool-use task, the model hallucinated content that did not originate from the actually retrieved tool result. The system capped the P2 score via hallucination cap. For content-critical tasks such as research, situation reports, or fact-bound summaries, this is a serious warning signal. An agentic model may be confident when working with tools. It may not, however, treat a fabricated finding as equivalent to a measured one.

This is not a minor criticism. Agentic models are purchased precisely because they are supposed to orchestrate tools. When the last mile — clean binding to actual tool results — becomes unreliable, productivity tips over into false precision. MiniMax M3 is capable here, but not blindly trustworthy. For operational DevOps tasks with human oversight, that is acceptable. For automated fact pipelines, it is a red line.

Reasoning and Logic: Correct, Sober, Not in Love with Its Own Mind

The reasoning module reveals a pleasant trait of this model: it thinks solidly without theatrical scaffolding. MiniMax M3 solves the classic guards puzzle correctly, with clean structure and in German. It walks through case distinctions clearly, even offers alternative formulations, and meets the format requirements precisely.

What is missing is not competence but intellectual surplus. The Judge rightly notes that the model does not surface the underlying double inversion as a conceptual core. MiniMax M3 solves the puzzle like a good technician, not like a mathematician with a love of elegance. That is not a criticism, merely a question of character. Those seeking exploratory deep reasoning will find less brilliance here than with models that put their thinking more on display. Those who simply want the correct answer get exactly that.

Notably, this result comes in standard mode. For a Thinking-Optional model, that is a good sign. It shows that the base configuration is already solid. You can also sense that reserves would be conceivable if an explicit thinking mode were available and activated.

Content Transformation and UX Proximity: Strong at Restructuring, Occasionally Too Sparse on Polish

In the Content Transformation module, MiniMax M3 delivers one of the most convincing performances in the entire report. The conversion into a German YouTube tutorial script succeeds not just formally but practically. Timestamps, spoken language, screen instructions, music cues, pattern interrupt, call to action, and even a cleanly integrated Easter egg: this is not merely correct — it is production-ready. You can tell the model has workflows, not just texts, in view.

At the same time, another qualitative sample reveals a weakness worth knowing. In a task involving inclusive German HR language, the output is noticeably too brief and stylistically too functional. The core work is sound, but the linguistic warmth is absent. Instead of an inviting formulation, the model delivers something closer to a directive. It captures the meaning but misses some of the cultural register of the text. MiniMax M3 is very good at restructuring. It does not always refine with the same care.

For German-speaking users, this matters. Good language models rarely fail here on grammar — they fail on tone. MiniMax M3 is competent in German, but not consistently sensitive. Where copy needs to create impact rather than merely convey information, you will need to edit more often than with its best technical outputs.

Documentation, Culture, and Language: Reliable, but Not Always Elegant

Documentation performance is strong overall. MiniMax M3 can organize information, present technical content accessibly, and write in a structured manner. This fits well with its large context window of one million tokens. Especially with long materials, branching documents, or multi-layered tasks, this capability is more than a number on a spec sheet — it is a practical working promise.

In the area of Cultural Intelligence, the verdict is somewhat mixed. The model remains linguistically clean and professional. It understands the broad requirements of culturally sensitive reformulations. But in finer German conventions, it occasionally comes across like a translator who has done very good work yet does not quite live in the target milieu. The HR reformulation example illustrates this clearly: inclusive yes, idiomatically appealing only in part. For business texts, that is often sufficient. For brand-sensitive communication, it is only a first draft.

API Cost Profile

MiniMax M3 is priced affordably, but not consistently economical in its output behavior. Precisely because this is a cloud model via OpenRouter, that is not an academic footnote but a cost question.

In the CLI domain, the model produces an average of 791 tokens against a fleet median of 312. That corresponds to 2.54 times the average across all tested models. In Code Quality, 5,438 tokens face a median of 2,921, putting it at 1.86 times. Content Transformation comes in at 3,620 versus 1,861 tokens, or 1.95 times. Cultural Intelligence stands out with 952 versus 290 tokens, a factor of 3.28. UX Writing also lands at nearly double, with 3,144 versus 1,577 tokens.

The good news: the model does not burn these tokens pointlessly. The bad news: it is often more verbose than economic necessity would require. Anyone deploying MiniMax M3 in a high-frequency API environment is buying not just quality, but a generous volume of words along with it.

Data Privacy and Data Sovereignty

For European organizations, MiniMax M3 is not a model for carefree everyday use from a data protection standpoint. The available cards indicate a calculated Sovereign Risk of HIGH. This has two dimensions. First, the model originates from MiniMax in Shanghai, China. Second, according to the Vendor Card, the provider is subject to Chinese law — specifically PIPL/CSL/DSL. For German and European users, this represents a clear third-country transfer risk with no EU adequacy decision in place.

The documentation lists China plus global MiniMax/OpenRouter partner routes as the data location. That is practical, but uncomfortable from a compliance perspective, since the actual data path can vary depending on routing. A GDPR DPA is listed as not available on the card. For organizations that must handle personal or confidential data in a GDPR-compliant manner, this is not a cosmetic flaw but a concrete deployment barrier.

The retention period adds another concern. It is listed as -1 days — meaning it is not reliably or transparently documented. Anyone wishing to deploy MiniMax M3 in production should restrict use to non-personal, heavily minimized, or already sanitized data. The fact that the weights are openly available does not automatically mitigate this point in this specific cloud scenario. The provenance of the weights is also rated HIGH risk on the card, given that origin and jurisdiction carry significant weight from a sovereignty perspective.

Conclusion

MiniMax M3 is a remarkable Frontier model with a clearly recognizable character. As an agentic generalist with vision capabilities, MoE architecture, and 23 billion active parameters, it delivers a rare combination of technical punch, structured working style, and production-ready output quality. Particularly in Code Quality, CLI, and Content Transformation, it operates at a level that commands genuine respect. It feels like a model that wants not just to answer, but to get things done.

Its weaknesses, however, are not cosmetic. The sporadic API instability is manageable; the problematic tail latency is noticeable in everyday use. More serious is the documented hallucination error in tool use. For agentic systems, that is the one point where confidence can quickly turn into liability. In German, too, MiniMax M3 does not always achieve the last degree of elegance. It is often accurate, but not always refined.

The recommendation is therefore differentiated. For DevOps-adjacent assistance, security reviews, code analysis, structured transformation tasks, and long-context work, MiniMax M3 is a very strong choice — all the more so given its aggressively attractive pricing. For fact-critical tool-use pipelines, automated research outputs without human oversight, and GDPR-sensitive enterprise data, caution is not a virtue but an obligation. MiniMax M3 is no smoke and mirrors. But you should not take every last claim it makes at face value without verification.

This evaluation was generated automatically based on the benchmark data. Model used: GPT-5.4 by OpenAI. The raw data and full methodology are documented in the GitHub project.