Mistral Medium 3.5

Mistral Medium 3.5 is Mistral AI’s Frontier model focused on agentic workflows and coding. With 128 billion parameters, the model operates with a context window of 256,000 tokens and supports multimodal inputs for text and image. Available as an Open Weights model under a Modified MIT license for local use or via cloud API, from a European provider environment with GDPR compliance.

Mistral AI Version 3.5 Commercial use permitted Dense 128 B (128 B active) 256 K Context 12/2025 $1.5 / $7.5 per 1M

  • Open Weights
  • Server
  • Mistral AI
  • Text
  • Vision
  • Instruction-Tuned
  • Real-Time

Sovereign Risk: LOW Mistral AI is headquartered in France. As an EU company, the storage and access to the weights are not subject to US CLOUD Act provisions or Chinese security laws.

LLM Model Review

Created on · Instruction-Tuned

With an overall score of 74.19%, Mistral Medium 3.5 presents itself as a remarkably fast Frontier model from the Mistral API: a generalist Dense-128B all-rounder with an Instruct character, an agentic lean, and multimodal capabilities, tested in standard mode without a Thinking toggle (n/a). The Speed Profile badge “Real-Time DevOps Expert” fits surprisingly well: this model responds like a system that doesn’t need a running start — it prefers to get to work immediately. That’s precisely why its weaknesses stand out all the more sharply when precision, depth of reasoning, or tool-grounded factual accuracy are required. Sovereign Risk: LOW — Mistral AI is headquartered in France, processes data in the EU according to the vendor, and under its current structure is not subject to the US CLOUD Act.

Header Grades: Stability and Reliability

Metric Value Rating Analysis
Timeout Rate 0/49 Stable The model ran with absolute stability and reliability throughout testing.
P95 Response Time 14.5 s Consistent Very low tail latency, almost no outliers.

For cloud models, stability is not a secondary concern — it is a contractual component of the value proposition. Mistral Medium 3.5 delivers a clean pass here. No failures, no drifting response behavior, no unpleasant surprises at the long end of response times. This is critical for agentic workflows, because an intelligent plan is worth little if the endpoint sporadically drops into the void.

Architecture and Character: Generalist with a Work Ethic

The pre-assigned category General, Instruct, Agentic, Multimodal captures the essence fairly well. As a Generalist, Mistral Medium 3.5 must be measured against the full breadth of the benchmark — not just coding or tool use. As an Instruct model, one can expect direct, task-faithful responses rather than sprawling internal monologues. Classified as Agentic, it must handle planning, structuring, and robust workflows. And as a multimodal model, it carries capabilities that a text-only benchmark can only partially surface. This matters: this test primarily measures the model’s linguistic and workflow-oriented character, not the full scope of its image competency.

Technical context also requires a note on scale. Mistral Medium 3.5 is a Frontier model and simultaneously a Dense Transformer with 128.0 billion parameters, all of which are active. Unlike Mixture-of-Experts models, there is no nominal parameter count masking a smaller active capacity. Expectations may therefore be set high. Those competing at this tier are not evaluated on good intentions.

The fact that the test run is marked n/a is not a deficiency but a consequence of the setup: as a commercial cloud model, there is no Thinking toggle available. What was tested is the default behavior — exactly the variant a typical API user gets without any special configuration. That is fair and informative. Because Mistral Medium 3.5 derives its value less from a performative thinking pose than from swift, mostly accurate execution.

Performance and Cost Profile

The badge “Real-Time DevOps Expert” says more than any raw throughput figure. It describes a model suited for interactive technical work: follow-up questions in terminal contexts, quick reviews, operational assistance under time pressure. Mistral Medium 3.5 feels exactly like that. It is neither a leisurely essayist nor a contemplative brooder. It is a working model.

This becomes interesting in combination with pricing. At $1.50 per 1 million input tokens and $7.50 per 1 million output tokens, Mistral Medium 3.5 is not cheap in the discount sense, but it is reasonably positioned within the Frontier context. Primarily because it behaves in a token-economical manner. No module exceeds the expected verbosity range. On the contrary: across all reported areas, the model sits below the fleet median. That is a genuine practical advantage. When billing via API, you are purchasing not just intelligence but also text volume. Mistral Medium 3.5 rarely writes beyond its means.

This economy aligns well with the Instruct side of the model. It responds with a certain discipline. Not ascetic, but controlled. The impression is clear: this is a model that wants to complete tasks, not dazzle for the sake of dazzling.

Code Quality and Security: Good Eye, Not Enough Bite

In the Code Quality domain, Mistral Medium 3.5 lands at 74.76%. That is not a security disaster, but it is also not a performance that would prompt a security team to set down their coffee cups. The qualitative log shows a model that broadly identifies vulnerabilities and structures them with formal cleanliness. The required Markdown table is in place, the response stays in German, and the list of identified gaps is extensive. In particular, the five implicit vulnerabilities were correctly identified. That is more than checkbox competency.

The catch lies in the depth of assessment. Mistral Medium 3.5 identifies several severe risks but rates two of them too leniently: Loose API Comparison remains at “High” where the standard sets “Critical”, and an IDOR vulnerability in the profile update is also downgraded. For everyday reviews, this is manageable. For serious security communication, it is a problem. Identifying threats while failing to assess their exploit chain sharply enough produces a false sense of calm. Security often fails not at the blind spot, but at the wrong priority list.

A second point compounds this: the model delivers a good list but no convincing attack narrative. It names weaknesses atomically but shows too little of how individual gaps chain together into a real attack path. Especially with topics like Type Juggling or Auth-Bypass, a correct label is not enough. You have to explain why it burns. Mistral Medium 3.5 writes here like a clean auditor, not like an experienced incident responder.

That said: for structured initial analyses, security checklists, and pre-review stages, this is usable. For final risk assessments, a human should sharpen the output. The model sees a lot. It does not always judge hard enough.

CLI and Agentic Suitability: Strong in the Tool Space, Not Free of Fabrication

The CLI and tool profile is one of the reasons this model remains interesting in practice. The partial score of 90.67% in the CLI benchmark is strong and convincingly supports the Agentic classification. Mistral Medium 3.5 structures technical tasks cleanly, thinks in operational steps, and behaves in shell-adjacent scenarios like a model that anticipated real work. This is where its productive side shows most clearly.

Yet it is precisely in this domain that the sharpest red flag of the entire review appears. In the Tool Use area, two hallucination violations occurred — specifically in content-critical tasks of the research or fact-bound tool evaluation type. In both cases, the model generated content that did not originate from the retrieved tool output but was instead fabricated. The system applied a hallucination cap to the score. This is not a stylistic weakness or a minor misread. For content-critical workflows, it is disqualifying.

For an agentic model in particular, this carries significant weight. Those who use tools must treat their output with iron discipline. Otherwise, Tool Use becomes nothing more than theatrical makeup for hallucinations. Mistral Medium 3.5 is strong in operational flow, but not at every moment sufficiently deferential to the source. For DevOps, CLI assistance, and structured technical dialogue, this often does not matter. For research, reporting, and any form of evidence-bound synthesis, it is a warning signal.

The ToolUse score of 59.17% captures this ambivalence precisely. The model can integrate tools. It is simply not consistently trustworthy in allowing itself to be constrained by them.

Reasoning and Logic: Correct, but with Loops in Its Head

In the Logical Reasoning domain, Mistral Medium 3.5 achieves 69.57%. That is respectable, but not outstanding for a Frontier model of this scale. The qualitative log for the guard puzzle is instructive: the solution is correct. The logic is sound. Only the path to get there is unnecessarily long, repetitive, and pedagogically unrewarding.

This is a classic distinction between “being right” and “explaining well.” Mistral Medium 3.5 reaches the destination, but it meanders. It restates the same insight in slightly varied form rather than naming the underlying mechanism once with precision. The gold standard articulates the double inversion clearly and supports it with structure. Mistral Medium 3.5 circles the same idea as if it does not trust its own punchline.

For classification purposes, this matters. As an Instruct model, brevity and directness are expected. When it then becomes verbose in logic tasks of all places, this reflects not deeper cognitive effort but a lack of editorial self-control. There are no visible reasoning tokens. This suggests internal inference. The result is solid here, but not elegant.

Metacognition Compliance (Reasoning): The model refuses to use the explicitly requested <thought> tags in 3/5 metacog tests, citing a consistent policy statement. The reasoning content itself is partially correct — the score deduction results from format non-compliance, not from reasoning errors. For comparison: in the tag-free reasoning_5* tests, the model achieves an average score of approximately 69.6%, which is on par with other models. CrucibleMark deliberately evaluates native zero-shot instruction compliance as a real-world characteristic — this deduction is methodologically intentional.

This passage is not a formality. Those building agentic chains depend on format compliance. When a model repeatedly overrides explicit markup requirements for policy reasons, that is an integration problem in practice. Not philosophically interesting — practically inconvenient.

UX Writing, Content Transformation, and Language: Professional, but Not Always with Final Refinement

In UX Writing, Mistral Medium 3.5 stands at 76.59%. This fits the overall character. The model formulates clearly, cleanly, and without artificial pathos. It is not the model that turns microcopy into great literature. But it writes comprehensibly, goal-orientedly, and with a good feel for specifications. The Instruct DNA pays off here.

In Content Transformation as well, with 76.88%, it delivers solid work. The log for the German YouTube script task shows a model that takes formatting rules seriously, stays in German, cleanly embeds production notes, and produces a complete, usable script. It understands the form of the task. That is half the battle.

The other half is polish. The Judge rightly notes the missing pattern interrupt in the middle section, an Easter egg that is more decorative than interactive, and an overall runtime that is too short. The result would be usable in practice, but not maximally optimized for retention and community engagement. Put differently: Mistral Medium 3.5 builds a solid video script, but not the script of someone who has YouTube mechanics in their blood.

In the Cultural Intelligence module, with 75.32%, the model comes across as pleasantly matter-of-fact. The inclusive revision of the toxic job posting succeeds very well. Problematic language is removed, gender bias is cleanly neutralized, and the text remains professionally readable. The only minor flaw is stylistic. The output lacks a degree of rhetorical energy. That is not a real error — more a matter of temperament. Those looking to transform provocative rough drafts into respectable business language will find a reliable tool here.

Documentation Quality: The Actual Weak Point

At 67.07%, Documentation Quality falls noticeably behind the stronger modules. This aligns with the observations from Code and Reasoning: Mistral Medium 3.5 can structure, but cannot always provide sufficient depth. Good documentation demands more than correct points in the correct order. It demands context, prioritization, connectivity, and the ability to make complex things not just accurate but teachable.

That is precisely where the model loses ground. It often explains usably, but too briefly or too linearly. Where a very good documentation model surfaces alternatives, pitfalls, and implicit assumptions, Mistral Medium 3.5 remains functional. That suffices for internal notes, quick drafts, and technical first passes. For documentation with training value, revision is required.

Hallucinations: Not a Cosmetic Flaw, but a Question of Trust

The hallucinations in the tool context deserve their own section, because they sharpen the character portrait of the model. Mistral Medium 3.5 does not hallucinate pervasively and does not generally come across as fabulist. That is precisely why the two documented cases are relevant. They do not indicate a persistently loose relationship with the truth, but a situational overstepping of source boundaries — exactly where it is least needed.

This is the uncomfortable kind of error. Not loud, not absurd, but plausible enough to slip through. For developers using the model as a fast operator in DevOps or shell contexts, the risk remains manageable. For factual work with tool output as the primary source: trust only with guardrails.

Data Privacy and Data Sovereignty

Here, Mistral Medium 3.5 plays an advantage that in European enterprises sounds not merely appealing but reflects procurement reality. The vendor Mistral AI SAS is headquartered in Paris, France, the applicable law is EU (GDPR), the stated data location is within the EU, and a GDPR DPA is available. For German and European companies, this is considerably more favorable than US-based vendors, because under the current vendor structure no CLOUD Act exposure exists.

The standard API operates with 30 days of data retention according to the Vendor Card. That is not zero risk, but it is transparent and manageable in an enterprise context. Also notable is the calculated Sovereign Risk: LOW. Combined with a Weights Provenance Risk of LOW, the result is an overall clean sovereignty profile. Those who treat European data sovereignty not as a marketing slide but as a procurement criterion will find here a rare combination of capability and legal grounding.

Conclusion

Mistral Medium 3.5 is a fast, disciplined, and remarkably work-oriented Frontier model. Its strengths lie in direct task execution, CLI-adjacent technical assistance, clean language output, and a refreshingly economical token footprint. As a Generalist, it delivers broad, usable overall performance. As an Instruct model, it follows instructions cleanly in most cases. As an Agentic-leaning candidate, it convinces more through structure and operational utility than through theoretical elegance. And as a multimodal model, it inevitably remains only partially visible within this text-only benchmark.

The weaknesses are clearly defined: Documentation Quality falls short of the Frontier standard, security assessments are at times too lenient, and in the Tool Use context, hallucination resistance is not strong enough for blind trust. The model is therefore not a universal genius, but a serious working instrument with a recognizable profile. Those seeking fast, clean technical assistance from a European vendor cloud will find substantial value here. Those looking to automate source-bound factual work or safety-critical prioritization should not leave Mistral Medium 3.5 alone in the room.

This evaluation was generated automatically based on the benchmark data. Model used: GPT-5.4 by OpenAI. The raw data and full methodology are documented in the GitHub project.