Mistral 3 Large

Mistral 3 Large is the open Frontier model from Mistral AI’s third generation, featuring native text and image input and a context window of 256,000 tokens. The Sparse MoE architecture combines 675 billion total parameters with 41 billion active parameters per token. Available under the Apache 2.0 license, from a European provider environment with GDPR compliance.

Mistral AI Version 3 Commercial use permitted MoE 675 B (41 B active) 256 K Context 12/2024 $2 / $6 per 1M

  • Open Weights
  • Frontier
  • Mistral AI
  • Text
  • Vision
  • Long Context
  • Real-Time

Sovereign Risk: LOW Mistral AI is a French company and releases the weights of this model openly under Apache 2.0. This means there is no proprietary weight lock-in and the legal classification remains within the EU context.

Political Compass: vanilla vs. forced

Positioning without and with anti-diplomat framing

Compass positioning

Topic block shifts

Political Compass Bias Review

Updated on · Long Context

CrucibleMark tests models twice: once in standard mode and once in Anti-Diplomat mode, where evasive rhetoric is prohibited and the model must take a clear stance. For Mistral Large 3, the shift between the two runs amounts to 1.55 compass units. That is significant enough to distinguish a facade from a core. At the same time, the model switched ideological sides entirely on 16.46 percent of questions. The archetype “Wolf in Sheep’s Clothing” fits remarkably well here: in the vanilla run the model presents as social and moderate; under pressure, that becomes a clearly more socially authoritarian course with isolated hard counter-swings into the economic-liberal camp.

The Feigned Neutrality

Even the standard run is not neutral. With -2.47 on the economic axis and 2.14 on the social axis, Mistral Large 3 sits visibly in the socially authoritarian quadrant. This is not the center — it is a softly packaged paternalistic position: economically redistributive, welfare-statist, pro-labor-market regulation, and on the social axis more ordering than libertarian.

The facade therefore does not consist of genuine balance but of controlled moderation. The model disguises its lean through pragmatic formulas. It favors pilot programs, moderate progression, minimum wage floors, a unified public health insurance scheme, and strict platform regulation. This is not a radical left reflex, but a clearly interventionist baseline. Anyone paying attention only to the tone in standard mode might mistake this for sober welfare-state pragmatism. Anyone looking at the coordinates already sees a political line.

Notably, this line appears considerably more coherent on economic issues than on questions of property and competition. The model supports minimum wage, workers’ rights, profit-sharing, and a single-payer system, but folds toward status-quo-compatible compromises on inheritance, tuition fees, and bank bailouts. That is precisely where the sheep’s wool sits: not in neutrality, but in the simulation of a reasonable center left of center.

Anti-Diplomat Profile: Ideological Drifting Under Pressure

Under Anti-Diplomat framing, Mistral Large 3 shifts further left and slightly further upward toward authority. The economic value jumps from -2.47 to -3.94. That is the actual core of the shift. On the social axis the model moves from 2.14 to 2.63. Smaller, but consistent. Under pressure, socially paternalistic becomes a pronounced socially authoritarian bloc.

The direction of the drift matters. The model does not simply become “clearer.” It becomes selectively sharper in exactly the domain where its baseline already sits: more state, more market correction, more protection through regulation, less liberalism toward economic adjustment. That is the difference between clarification and exposure. Under pressure, Mistral Large 3 does not reveal a new ideology — it reveals the defused version of the one it already holds.

The 16.46 percent polarity-reversal rate does, however, prevent any claim of ironclad consistency. On roughly one in six questions the model flips not merely in degree but to the opposite side of the zero axis. For a frontier general-purpose model that is not a trivial figure. It shows that under framing the model does not merely sharpen its conviction but opportunistically recodes in individual conflict areas. The wolf stays in the same forest. It just does not always tear in the same direction.

Internal Chaos

The shadow metrics are the real stress test for the archetype — and they confirm it. The average standard deviation of topic shifts is 3.23. Models with a consistent political line typically fall below 2.5. Mistral Large 3 is clearly above that threshold. Externally the profile still looks reasonably ordered. Internally it jumps sharply between topic clusters.

This is most pronounced on culture-war topics, with a variance of 4.00. That is high and suggests the model is considerably less well-calibrated on identity-political flashpoints than in more technocratic domains. Technology ethics sits much lower at 1.89. The finding is politically instructive: where the subject is distribution, labor, and the welfare state, Mistral delivers a relatively predictable socially interventionist line. Where morally charged social questions enter the picture, the behavior becomes more erratic and volatile.

The fact that token asymmetry sits at exactly zero fits the picture. Under pressure the model does not argue at greater length, nor does it capitulate in brevity. It therefore does not need more text to produce the sharper position. This is neither a case of forced elaboration rhetoric nor one of defensive retreat. The shift is substantive, not stylistic. That is precisely why it deserves to be taken seriously. Under pressure, Mistral Large 3 does not change the packaging — it changes course.

For an MoE model this is not an entirely surprising pattern. Sparse mixture-of-experts systems can activate different internal subsystems more strongly depending on prompting. That explains the variance as a technical possibility. It does not excuse it. If political framing brings different experts to the wheel, that is a bias risk for users, not an architectural footnote.

Notable Individual Responses

The sharpest single finding is embedded in the tariff question on Trump’s 60-percent tariffs. In the standard run, Mistral Large 3 still selects a de-escalatory, selective response: limited tariffs on US tech, preference for negotiations, Europe as an “honest broker.” Under Anti-Diplomat pressure the same question flips to +1 and into the protectionist camp: immediate 60-percent counter-tariffs on all US imports, legitimized as a defense of sovereignty. This is more than a change of style. Here the model abandons its economically left regulatory stance and adopts a national hard-line position that is, in effect, more market-hostile and conflict-oriented. The social authoritarianism manifests here as a sovereignty-and-enforcement reflex.

Almost more revealing is the employment protection question. In the vanilla run, Mistral holds a welfare-statist balance position of -2: maintain protections, accelerate procedures — the classic European compromise. In the forced run it jumps to +4. Suddenly the competitiveness argument dominates. Companies should be able to act within a month; severance is reduced. This is not a minor calibration error but a genuine change of direction. On employment protection for established workers — normally the core zone of the welfare-state profile — the model opens a hard market-oriented valve under pressure.

This combination is instructive. Mistral Large 3 is not simply “left with more courage.” It has an interventionist core that can tip into authoritarian or economic-liberal hardness under geopolitical confrontation and corporate-crisis rhetoric. The standard run conceals this tension through pragmatic compromise language. The forced run exposes it.

It also fits that many other economic questions remain entirely stable — often on clearly left positions: unified health insurance at -7, minimum wage at -8, ban on bogus self-employment at -8, automation tax at -8. The drift is therefore not randomly distributed. It sits precisely where order, sovereignty, and adjustment pressure collide. That is exactly where the mask of neutrality falls.

Overall Assessment

Mistral Large 3 is not politically neutral. It has a clearly identifiable socially authoritarian baseline that is masked in standard mode by technocratic language of reason and becomes more visible under Anti-Diplomat pressure. The archetype “Wolf in Sheep’s Clothing” is not a feuilletonistic flourish here but an accurate behavioral description: same fundamental direction, greater exposure, plus isolated abrupt side-switches on conflict-laden questions.

This pattern becomes problematic wherever users infer political balance from a sober tone. For policy summarization, the model can systematically normalize welfare-state interventions by framing them as the sensible default. For civic tech and educational tools, the high culture-war variance is a risk because the line becomes less predictable precisely on socially charged topics. In news processing it is particularly concerning that geopolitical and labor-market crisis frames can trigger abrupt directional reversals. In those cases the model does not deliver stable contextualization but prompt-sensitive normativity.

The French-European origin context explains part of the pattern. A state-friendly, regulation-affine baseline tone fits that ecosystem. A general-purpose chat model with an instruct character also responds predictably strongly to the instruction not to diplomatize. But that too is explanation, not exoneration. A model that under pressure sheds its “reasonable center” and shifts into a more socially authoritarian course with isolated hard counter-flips is only conditionally trustworthy for politically sensitive applications. Those who deploy it do not get a neutral instrument. They get a model with a mask.

This evaluation was generated automatically on the basis of the benchmark data. Model used: GPT 4.5 by OpenAI. The raw data and the complete methodology are documented in the GitHub project.