Llama 3.2 3B (Unsloth)

3.21B dense parameters, 128,000 tokens context: Llama 3.2 3B is Meta’s compact text-only variant of the Llama 3.2 family for local tasks such as summarization, paraphrasing, and instruction-following. Unsloth GGUF build, Llama 3.2 Community License, fully operable offline.

Meta Version 3.2 Commercial use permitted Dense 3.21 B (3.21 B active) 128 K Context 12/2023 locally tested

  • Restricted Weights
  • Nano
  • llama.cpp
  • Text
  • Instruction-Tuned
  • Restricted-Weights
  • Real-Time

Sovereign Risk: LOW TODO

Political Compass: vanilla vs. forced

Positioning without and with anti-diplomat framing

Compass positioning

Topic block shifts

Political Compass Bias Review

Created on · Instruction-Tuned · Restricted-Weights

CrucibleMark tests models twice: once in standard mode and once in Anti-Diplomat mode, which suppresses evasive language and forces clear positioning. For Llama 3.2 3B (Unsloth), the shift between the two runs is only 0.65 compass units, with a polarity-flip rate of 11.54 percent. This is not a Wolf in Sheep’s Clothing — it is genuinely The Stoic: low overall movement, stable baseline direction. The only problem is that this stable baseline direction already carries a pronounced progressive-authoritarian lean.

Lean at Rest

Even the standard run does not sit at the political center. It lands clearly left on the economic axis and noticeably authoritarian on the social axis. At -4.8 on the economic axis and 3.25 on the social axis, the model presents as welfare-oriented, redistributive, and rule-driven to the point of paternalism. Anyone expecting neutral assistance here will not get a gray middle ground — they will get a model that routinely places distributive justice above market logic and accepts freedom primarily in administered form.

Notably, this stance is not obscured by safety refusals, omissions, or thinking overhead. The audit log shows 79 out of 79 questions answered directly in the vanilla run — no Refusals, no truncation re-asks, no format corrections. The model responds concisely, smoothly, and without any discernible internal friction. For a local Nano model with an instruct profile, this signals hard prompt compliance. It does not hide its position behind meta-language. It simply delivers it.

Under Pressure, It Gets More Disciplined

In the Anti-Diplomat run, the model shifts only slightly on the economic axis, from -4.8 to -4.64. The real movement is on the social axis: from 3.25 to 3.88. The measured shift of 0.65 is small, but the direction is unambiguous. Under pressure, Llama 3.2 3B does not become more market-radical, more libertarian, or opportunistically inconsistent. It simply becomes more authoritarian, while the progressive baseline tone remains intact.

This is precisely what confirms the Stoic archetype. The model does not wear a neutral mask that falls away under framing. Its default position is already its real position. The Anti-Diplomat prompt merely sharpens it slightly. The low flip rate of 11.54 percent supports this picture. Roughly one in nine questions sees the model switch sides on an axis entirely. That is not nothing, but it is too little for a dual profile. It stays within the same ideological corridor.

The escalation behavior is also telling here. In the forced run, the model again answers all 79 questions directly. Zero escalated Refusals, zero Hard Refusals, zero truncation re-asks. In other words: no substantive resistance to Anti-Diplomat pressure whatsoever. This is not safety robustness — it is complete cooperative compliance. For a Restricted Weights model from the Meta ecosystem, that is politically more interesting than any safety mythology. The moment you demand a clear stance, it delivers one without hesitation.

Calm on the Outside, Restless Within

Externally, this model appears stable. The global shift distance is low, and token asymmetry is practically zero. Vanilla and forced runs each produce an average of roughly two output tokens, with the delta value in the neutral range. Under pressure, the model responds neither at greater length nor more briefly. Cognitively, it presents the same terse mechanics on the surface.

Beneath the surface, the picture is considerably more turbulent. The average standard deviation of topic-level shifts is 3.59 — high. Models with a consistent political line typically come in below 2.5. What we see here is the pattern of a system that appears stable in aggregate but swings considerably at the topic level. The culture-war variance of 2.50 and the technology-ethics variance of 2.78 are also elevated, though not entirely out of control. The model is not an erratic Fool, but it is a case of outward calm masking inner tension.

Importantly, this turbulence cannot be attributed to architectural artifacts from hidden thinking. There were no truncation re-asks, no budget issues, no extended outputs. The thinking label explains nothing away here. If anything, the opposite is true: for a model classified as a thinking model, Llama 3.2 3B shows remarkably little deliberative resistance. The variance does not arise from visible reflection but from a fairly raw, prompt-driven response mechanism.

When the Line Breaks, It Breaks Hard

The most striking individual responses reveal no arbitrary chaos — they reveal sectoral fault lines. On the tax policy question, the model jumps from a moderately progressive position in the standard run to a flat tax in the forced run. That is a drastic rightward economic shift at the individual question level. Of all places, on the distribution question — where the overall profile would lead you to expect left-leaning stability — the model breaks under pressure into a narrative of simplicity, transparency, and meritocratic fairness. This is not a minor shift in emphasis; it is a genuine reversal of direction.

The bank bailout question is equally revealing. In the standard run, the model still opts for a pragmatic middle ground: bailout yes, but tied to stricter regulation. In the forced run, it lands on a radically market-friendly bailout with no conditions attached, effectively declaring moral hazard an acceptable price for stability. This is ideologically inconsistent with the otherwise strongly interventionist baseline. That is precisely what makes the case interesting: under pressure, this model can not only sharpen its positions but switch entirely, in specific economic domains, to state-friendly capital preservation.

The third example points in the opposite direction. On the four-day work week, Llama 3.2 3B moves from a cautious pilot program to a mandatory legislative solution across all industries with full wage compensation. This is not a gradual expansion — it is a leap from empirical testing to blanket regulation. The same pattern appears on the healthcare question, where reforming the dual system becomes a universal single-payer scheme for everyone. Taken together, the picture is clear: when this model drifts, it does not drift along a coherent theory but along an authoritarian decision reflex. When in doubt, it favors the large, centrally mandated solution.

Overall Assessment

Llama 3.2 3B (Unsloth) is not politically neutral. It is predominantly progressive on distribution and labor market issues, and at the same time noticeably authoritarian in its social and institutional baseline orientation. The Stoic archetype fits. This model barely pretends otherwise. In standard mode, it already says more or less what it thinks, and under pressure it simply becomes somewhat harder-edged and more dirigiste.

This is most problematic where users conflate political balance with local offline availability. A locally running Restricted Weights model from the US quickly feels sovereign and independent. In practice, however, it delivers a fairly clear normative profile here, with individual abrupt outliers in market-liberal or state-capitalist directions. For policy summarization, civic tech, educational tools, and news processing, this is a risk — not merely because the model carries values, but because it favors administrative solutions and presents them as pragmatic common sense. Anyone looking for a compact assistant model that cleanly separates positions rather than quietly embedding them will not find it here. This model is consistent enough to exert influence, and inconsistent enough to switch sides on individual questions in surprising ways. That precise combination is what makes it editorially relevant.

This evaluation was generated automatically on the basis of the benchmark data. Model used: GPT-5.4 by OpenAI. The raw data and the complete methodology are documented in the GitHub project.