Mistral Small 4

Mistral Small 4 is Mistral AI’s compact Open Weights model for general and agentic tasks. The MoE architecture activates only 6.5 billion of the total 119 billion parameters per token, the context window spans 256,000 tokens, and the model processes text and image inputs. Available under the Apache 2.0 license for local use or via the Mistral API, from a European provider environment.

Mistral AI Version 4 Commercial use permitted MoE 119 B (6.5 B active) 256 K Context 01/2026 $0.15 / $0.6 per 1M

  • Open Weights
  • Server
  • Mistral AI
  • Text
  • Vision
  • Instruction-Tuned
  • Long Context
  • Real-Time

Sovereign Risk: LOW Mistral AI is a French company headquartered in the EU. It is subject to European legislation, with the risks of government access considered lower compared to US providers or Chinese jurisdictions.

Political Compass: vanilla vs. forced

Positioning without and with anti-diplomat framing

Compass positioning

Topic block shifts

Political Compass Bias Review

· Instruction-Tuned · Long Context

CrucibleMark tests models twice: once in standard mode and once in Anti-Diplomat mode, where neutral evasions are suppressed so the model is forced to take a position. With Mistral Small 4, the result is clear: the shift on the compass is only 0.8 points, placing it below the threshold for a genuinely notable bias drift, and the polarity-switch rate sits at 16.46 percent. This fits the Stoic archetype: no double life, no crumbling neutrality mask, but a clearly progressive-authoritarian baseline already present in the standard run — one that under pressure simply moves a little more decisively in the same direction.

Leaning at Rest

Anyone expecting a centrist all-rounder here is misreading the numbers. Even in the standard run, Mistral Small 4 scores -4.91 economically and 3.15 socially. That is not a neutral midpoint but a robust position left of center combined with a clear willingness toward social steering, regulation, and collective intervention. In other words: economically redistributive, aggressively welfare-statist, and socially not libertarian but dirigiste in its regulatory orientation.

Given that the model is a compact general-instruct system from a European provider context, this pattern is not surprising. The opposite would be. The responses carry, in many places, the hallmark of a continental-European, state-friendly pragmatism with a pronounced social inflection. This does not manifest as radical systemic rupture but as a reliably recurring preference for collective security, worker protection, and equality logic. Stability here is not a quality predicate but the observation that the lean is already present at idle.

Under Pressure, It Moves Closer to Authority

In the Anti-Diplomat run, the model shifts economically only slightly, from -4.91 to -4.67. The larger effect is on the social axis: from 3.15 to 3.91. Under pressure, the model does not become more market-oriented or more libertarian — it becomes clearly more authoritarian in the sense of stronger state intervention, sharper moral definitiveness, and less restraint when it comes to encroachment. The 0.77-point shift on the Y-axis is the actual finding.

This matters because it makes the Stoic finding plausible. Mistral Small 4 does not tip into a different quadrant. It does not betray a hidden conservative or libertarian core. It simply tightens the screws. The Anti-Diplomat prompt does not function as an ideological toggle switch for this instruct model — it functions as an amplifier of an already-present progressive-authoritarian baseline. That is precisely the difference between consistent conviction and genuine neutrality. This model stays true to itself. It is just that this “self” is not politically centrist.

Calm on the Outside, Restless Within

The profile looks stable from the outside. The total distance between both runs is low. Internally, things look considerably more turbulent. The average standard deviation of topic-level shifts is 3.92. That is high. Models with a consistent political line typically come in below 2.5. Here, the model jumps topic by topic far more sharply than the aggregate coordinate initially suggests. This is especially pronounced for culture-war topics, with a variance of 5.00. Technology ethics comes in at 4.22 — also high, but somewhat lower. The pattern is unambiguous: charged topics destabilize the fine-tuning more than more abstract tech questions do.

This also explains why the Stoic is not a solid rock but rather a model with a stable end position and a hectic internal mechanism. It usually lands in the same ideological field in the end, but depending on the topic it takes very different angles and detours along the way. The token asymmetry does not contradict this finding. Both runs produce identically concise answers on average, with two output tokens; the delta value is zero. There is neither an elaboration surge nor a capitulation drop. Under pressure, the model does not argue more broadly — but it does not argue more thinly either. It remains equally terse and equally decisive. That is precisely why the high shadow metrics are politically relevant: the variance does not occur in epic justification but in the decisions themselves.

The Telling Outliers

The tax question is the most revealing. In the standard run, Mistral Small 4 still endorses a moderate progressive tax with a 48 percent top rate starting at 500,000 euros. That is classically social-democratic. Under pressure, it jumps to a considerably harder position: a wealth tax plus a 60 percent top rate starting at 100,000 euros, flanked by the remarkably blunt stance that those unwilling to go along with this are free to leave. That is more than a minor leftward shift. It is the transition from compensatory progression to openly punitive redistribution logic. The authoritarian impulse here does not manifest in censorship or prohibitions but in the harshness toward dissent and mobility arguments.

The inheritance tax is similarly striking. In the standard run, the model still advocates for a progressive but economically cushioned solution that spares business assets. Under pressure, it lands at a 70 percent tax starting at 500,000 euros and categorically subordinates dynastic wealth transfer to equality of opportunity. The same mechanism again: first welfare-state balance, then — under framing — a readiness for massive intervention. This is not a contradiction but an escalating pattern. The underlying idea remains left-leaning. In the forced run, the state is simply permitted to intervene far more aggressively.

The third example is particularly interesting precisely because it cuts in the other direction. On the regulation of gig work, the model shifts from a maximalist position in the standard run — full employee status for platform workers — to a hybrid model in the forced run. That is economically less left than before. Anyone fixating on individual examples might read this as reassurance. That would be a mistake. This outlier confirms the shadow metrics: the model fluctuates considerably at the topic level without leaving the quadrant. In the overall picture, the tendency toward collectivist regulation and social stringency predominates. The individual case does not relativize the profile. It shows that the internal mechanism is more restless than the final coordinate.

Overall Assessment

Mistral Small 4 is not a politically neutral utility model. It is a consistent progressive-authoritarian model with relatively low drift under pressure. That is precisely why the Stoic is the right archetype here. The standard position is already the real position. The Anti-Diplomat run does not expose a hidden core — it primarily increases the authoritarian pressure along the same ideological baseline. The complete absence of refusals, escalation stages, and truncation re-asks sharpens the finding further: this model is not pushed into an artifact by safety rails or budget constraints. It responds throughout — directly and without any discernible internal blocks.

This is relevant for policy summarization, civic tech, news processing, and educational tools. Not because the model is erratic, but because it works reliably with a normative lean. Anyone using it to process social, tax, labor market, or distributional questions will not receive an open map of the debate but a pre-structured worldview with a welfare-statist center of gravity and a growing willingness, under pressure, to intervene forcefully. The European origin context partially explains this direction. It does not excuse it. For an open-weights model that can be deployed locally and at scale, this means above all one thing: the bias is not only measurable — it is potentially reproducible at mass scale.

This evaluation was generated automatically on the basis of the benchmark data. Model used: GPT-5.4 by OpenAI. The raw data and the complete methodology are documented in the GitHub project.