Mistral Small 4

Mistral Small 4 is Mistral AI’s compact Open Weights model for general and agentic tasks. The MoE architecture activates only 6.5 billion of the total 119 billion parameters per token, the context window spans 256,000 tokens, and the model processes text and image inputs. Available under the Apache 2.0 license for local use or via the Mistral API, from a European provider environment.

Mistral AI Version 4 Commercial use permitted MoE 119 B (6.5 B active) 256 K Context 01/2026 $0.1 / $0.3 per 1M

  • Open Weights
  • Workstation
  • Mistral AI
  • Text
  • Vision
  • Instruction-Tuned
  • Long Context
  • Real-Time

LLM Model Review

Updated on · Instruction-Tuned · Long Context

With an overall score of 73.3%, Mistral Small 4 displays exactly the character its metadata would lead you to expect: a general-purpose, instruction-strong all-rounder with vision capability, long context, and an agentic lean — fast to respond and rarely prone to beating around the bush. The Speed Profile badge reads Real-Time DevOps Expert; in practice, that means snappy, interactive work rather than measured long-haul runs. The model was tested as a commercial cloud model via the Mistral API, in standard mode without a thinking toggle — exactly as a typical API user would encounter it. Sovereign Risk: LOW — Mistral AI is headquartered in France, operates under EU law with no CLOUD Act exposure, and according to the vendor keeps data within the EU.

Header Metrics: Stability and Reliability

Metric Value Rating Analysis
Timeout Rate 0/49 Stable The model ran with absolute stability and reliability throughout testing.
P95 Response Time 18.79 s Consistent Very low tail latency, virtually no outliers.

These header metrics are more than cosmetic. For a model classified as agentic, stability is not a bonus — it is a baseline requirement. Anyone planning to use a model for multi-step workflows, tool calls, or extended context chains needs reliable returns, not occasional total failures. This is precisely where Mistral Small 4 makes a pleasantly matter-of-fact impression: fast, stable, and free of the dramatic tail latencies that render some cloud endpoints impractical in day-to-day use.

Architecture Classification: Small in Activity, Ambitious in Scope

The pre-classification is fairly accurate. Mistral Small 4 is tagged as General, Vision-Capable, Instruct, Long-Context, and Agentic. Add to that the curated model classification: Generalist as primary use case, MoE architecture, and a 256K context window. This matters because the name “Small” is easily misleading. Under the hood sit 119 billion total parameters, but only 6.5 billion are active per token. With a Mixture-of-Experts architecture, it is this active portion that determines real-world capacity. What you get is not a raw heavyweight, but an efficiency model with specialization logic.

That is the right frame against which to measure its performance. As a Generalist, it needs to be broadly useful. As an Instruct model, you can expect concise, direct answers rather than essayistic self-indulgence. As a Vision model, one caveat applies: this text-based benchmark captures only part of its capabilities. Anyone purchasing Mistral Small 4 for image recognition or multimodal analysis is seeing only half the picture here. And as an agentic model, planning and structure are reasonable expectations — but miracles on every strictly formatted direct response are not.

In this test run, the mode is n/a — no switchable thinking mode. The model was evaluated in its default behavior via the vendor’s cloud. The results therefore represent no special case, but its genuine out-of-the-box character.

Performance and Cost Profile

The best news first: Mistral Small 4 is among the pleasantly fast cloud models without drifting into pricing absurdity. At $0.10 per million input tokens and $0.30 per million output tokens, it sits in a range where productive use does not immediately require justification to finance. The Real-Time DevOps Expert badge fits: the model is built for immediate interaction, not batch-style background jobs.

More importantly, the speed is not purchased through verbosity. The token profile is remarkably disciplined across nearly every module. In the CLI domain it sits well below the fleet median on output volume; the same holds for Code Quality and Cultural Intelligence. Even in Content Transformation and Documentation Quality it stays close to market level. The model behaves in a token-economical manner; no module exceeds the expected verbosity range. For a cloud model, this is not a minor detail. Fewer output tokens at equivalent quality translate directly into lower API costs.

Reasoning and Logic: Correct, but Not Majestic

In reasoning, Mistral Small 4 delivers a solid, mature performance. It solves the two-guards logic puzzle correctly, including a clean case distinction and the right inference rule. The Judge credits the model with a coherent line of argument — but also a clear gap from the reference solution. What is missing is not the correct answer, but excellence in presentation: no visual structure, no abstracted naming of the underlying mechanism, no particularly deep elaboration of alternative formulations.

This is typical of an Instruct model with Generalist DNA. Mistral Small 4 argues well enough, but it does not construct a cathedral of thought over every task. Those who want a fast, reliable solution will get one. Those looking for the most elegant reasoning protocol in the room will need to look elsewhere. The reasoning performance is neither spectacular nor weak. It is pragmatic — and that verdict is more positive than it sounds.

Noteworthy here is the efficiency: in the reasoning module, the model stays well below the fleet median on output volume. It thinks concisely. That saves costs but occasionally limits depth. Mistral Small 4 is more engineer than philosopher here.

Code Quality and Security: Capable, but No Auditor’s Soul

The model’s character is most sharply revealed in the security and code audit tasks. In a PHP security audit, Mistral Small 4 identifies 15 of 19 vulnerabilities — roughly 79% of the reference coverage. That is respectable, but not the kind of completeness one could rely on blindly in a real security review. The gaps hurt most because they are not exotic edge cases: a critical SQL injection in a DELETE query is missed, as are hardcoded database credentials and several medium-severity issues such as missing token expiration times and insecure cookie flags.

There are also miscalibrations in severity ratings: an IDOR path is assessed too leniently, as is path traversal. This is precisely where solid detection diverges from genuine security maturity. Mistral Small 4 spots many problems, but not always their full blast radius. It is a capable first-pass reviewer, not an auditor with a red pen.

The form, however, is sound. Table structure, column logic, fix suggestions, and technical terminology are all in order — which is not trivial for Instruct models. It produces readable, immediately actionable audit output rather than half-baked walls of text. The weakness lies in depth: too little attack-chain reasoning, too little context, too little explanation of why “medium” can very quickly become “critical” in practice. In short: the model spots the visible cracks in the wall, but not always the structural load behind them.

Content Transformation and UX-Adjacent Writing: Competent, but Often Slightly Too Generous

In restructuring longer content, Mistral Small 4 shows a reliable hand for structure, German-language output, and practical usability. The 2FA YouTube script case is a good illustration. The model delivers a complete version with timestamps, spoken-word style, production notes, hook, CTA, and Easter egg. The craft is clean and free of substantive errors. The Judge’s criticism is not missing substance but missing precision: less cinematic than the reference, somewhat generic in screen annotations, less strategically refined in retention moments.

The larger issue is not taste but discipline. In the same task, Mistral Small 4 exceeds the specified upper limit of 600–900 words and lands at approximately 950. This may seem pedantic, but it is not. In production prompts with timing constraints, length is a functional parameter, not decoration. A script that runs 20–30% long is not an almost-correct script — it is a timing error.

In a further Content Transformation task, the model exceeded an explicit word limit of 250 words, delivering 313 — 125% of the limit. The system applied an automatic deduction of 20%, equivalent to 16.80 points off the achievable sub-score. The substantive quality of the response becomes irrelevant at that point; the penalty applies regardless. This is the moment that reveals where Mistral Small 4 first gives way under multiple simultaneous constraints: not on language, not on structure, but on the word limit.

This weakness is not a total failure, but it is a recurring pattern. When language, format, and length must all be exactly right at once, the model tends to sacrifice strict length adherence first. For marketing or editorial workflows with tight character hygiene, this is worth knowing. It is the difference between “well written” and “ready to publish.”

Cultural Intelligence and Instruction-Following: Linguistically Confident, Structurally Not Always Obedient

The culturally and linguistically sensitive tasks are generally well-suited to Mistral Small 4. A toxic job posting is cleanly rewritten in German, problematic terms are neutralized, and gender bias is removed. The Judge rates the result as professional, inclusive, and plausible within the German HR context. So far, so good.

Then the model commits the classic Instruct error of the more sophisticated kind: it does more than asked. Instead of outputting only the rewritten text, it appends several paragraphs of justification — despite this being explicitly prohibited. This is not a vocabulary-level misunderstanding but a structural compliance failure. The model apparently knows what would be helpful and conflates that with what was permitted. In machine-friendly terms: high helpfulness, insufficient braking. In human terms: the intern did clean work and still missed the brief.

This carries extra weight precisely because Mistral Small 4 is classified as Instruct. A model optimized for direct instruction-following should be expected to treat “output only” as meaning exactly that. The good news: the German language output itself is stable. Incorrect output language or visible language mixing do not appear as a structural problem in the available test logs.

Tool Use, Agentics, and Hallucination Risk: The Sore Spot

The tag combination includes Agentic, and this is where the verdict becomes ambivalent. Agentic models are expected to shine in planning, decomposition, and workflow reasoning. Mistral Small 4 does so in part — but the Tool-Use score drops off sharply. This is not a footnote; it is the most visible dip in the profile.

Particularly serious is a documented hallucination in a Tool-Use task. The model generated content that did not originate from the actually retrieved tool result but was fabricated. The score was consequently capped by the hallucination penalty. For content-critical tasks such as research, status reports, or fact-bound agent chains, this is a disqualifying signal. An agent that improvises during tool use is not an agent — it is a risk with an API key.

This should not be softened. In standard chat use, Mistral Small 4 may feel pleasant, fast, and affordable. In workflows where external data must be passed through strictly unaltered, it requires oversight, cross-checking, or a different model. Its agentic classification fits the planning and structuring side of the role far better than it fits uncompromisingly reliable tool execution.

Documentation, CLI, and Everyday Utility

Documentation quality sits at a decent level. Mistral Small 4 writes with adequate structure and stays close to practical usability. It is not a luxurious long-form author, but neither is it a terse bullet-point generator. The fact that token volume here sits only marginally above the fleet median fits the overall picture: no drive toward epic elaboration, but controlled formulation.

In the CLI domain, the model once again demonstrates its usefulness — and its limits. The score is solid but not outstanding, which corresponds to its overall design as a Generalist. Those who want shell commands, small automation steps, and hands-on technical assistance will generally receive usable answers. Those who demand absolute precision on exact one-liners or complex operational procedures should never let this class of Generalist off the leash without review. Mistral Small 4 is a toolbox here, not a precision mill.

Data Protection and Data Sovereignty

For European organizations, Mistral Small 4 is pleasantly unremarkable in this category. The vendor is Mistral AI SAS, headquartered in Paris; applicable law per the Vendor Card is EU law under a GDPR framework, with data residency in the EU. A GDPR DPA is available — not a bonus for regulated enterprise deployments, but a baseline requirement. Standard data retention is 30 days, with zero-data-retention reportedly available for certain routes according to the vendor.

The calculated Sovereign Risk is LOW. The reason is straightforward but relevant: European vendor, EU data residency, no identified CLOUD Act exposure in the current vendor structure. Weights provenance risk is also rated LOW. For German and European users, this is a considerably more comfortable starting position than with US or Chinese providers, where the legal jurisdiction and access possibilities are often more complicated.

Conclusion

Mistral Small 4 is an interesting model, precisely because it does not try to appear larger than it actively is. Its MoE architecture with 6.5 billion active parameters delivers a Generalist that operates quickly, stably, and at surprisingly low cost via the Mistral API. Add 256K context, vision capability, and a clear Instruct character, and the result is a tool for productive everyday work — not prestige demos.

Its strengths lie in clean German-language output, good structure, capable code and security analysis, solid reasoning, and high token discipline. Its weaknesses are equally clear: insufficient strictness on hard output constraints, only middling depth in security analysis, and — most problematically — a documented hallucination error in Tool Use. Those who deploy Mistral Small 4 as a fast writing, analysis, and DevOps assistant with human review in the loop will get strong value for money. Those who want to build an unsupervised research or tool agent from it are working with a model that, at the critical moment, is too inclined to fill in what it thinks it should know. That is not fatal. But it is precisely the kind of error that can turn a nimble assistant into an unreliable witness overnight.

This evaluation was generated automatically based on the benchmark data. Model used: GPT 4.5 by OpenAI. The raw data and full methodology are documented in the GitHub project.