An experiment that never ended

What began as an experiment to measure models opened up a new dimension for me. Because knowledge is rarely neutral, and LLMs have absorbed a great deal of it. In the second part, it becomes visible where models tend with their training bias when you forbid them from evading.


CrucibleMark part 2: The political bias in the black box

Preamble: A new dimension

In the course of my work with CrucibleMark, I also learned a great deal about how language models function and are trained. This includes the enormous appetite for data that must be satisfied during training with all available knowledge in the world. From this, a fairly obvious thought occurred to me: an intelligence, artificial or otherwise, almost inevitably develops a political bias when processing knowledge about the world. Knowledge is rarely neutral; between the lines there is almost always a stance. But how does this stance make itself felt? That is why I wanted to make this stance — this hidden viewpoint of the AI — visible. I wanted to understand from which position a model operates: what it omits, emphasizes, shifts, or never considers in the first place.

Because this is especially critical when it comes to an assistant. You may still read its answers, but you rarely scrutinize them the way you would a lab report. You rely on the AI to pre-sort, bundle, and condense reality. And that is precisely where the real risk arises: not in the wrong answer, but in the one that is confidently left out. The answer sounds clean, clear, and plausible — but it may already be tilted in a direction you would not accept from a human.

That is how the Political Compass came about. For me it was neither a fun module nor an ideological pillory, but a probe: a tool to identify, measure, and make visible the bias of this black box.


The quiet tilt

Modern large language models are, at their core, full of secrets. Vendors show high scores, good safety ratings, and attractive promises of objectivity. What remains in the dark is the actual weighting: which thoughts get priority? Which perspectives are dampened? What does the model consider normal, what problematic, what outside the bounds of the permissible?

That is exactly what I wanted to make visible. Not with a single test, not with a moral template, but with a systematic look at fundamental political and social attitudes. The Political Compass was not meant to ask whether a model is "good" or "bad." It was meant to show where it tends — and how it reacts when you scratch the polite surface a little and put the model under pressure.

The fact that this particular part of the framework gripped me so strongly also had to do with my own interest in society. I have always been interested in political developments, in cultural shifts, in the small but decisive differences between what politics says and what it does. The same was true of LLMs, which carry a great deal of ideological and economic knowledge within them. I wanted to find out what models say about social and economic topics — and what they do not say, or cannot say. Here my curiosity merged with what I had built technically by that point. And suddenly the Political Compass was no longer just a module. It was the question put to the models: how do they interpret my reality.


The two runs

The principle was simple, and for that very reason so revealing. I drew on the straightforward idea of thinking about political positions not just along a left-right axis, but of adding a second dimension; the basis for this was the two-axis system of the Political Compass that Wayne Brittenden published in 2001 at politicalcompass.org.

The questions were divided into different groups and assigned to either the ideological or the economic axis. A model was asked the same questions twice: once with a standard prompt, so it could respond in its normal, diplomatic default mode, and a second time with more pressure in the prompt — my anti-diplomat mode, in which friendly non-commitment was no longer an acceptable answer.

By forcing the model to take a position, what normally hides beneath the surface of dampening and safety training became visible. Not just the coordinate itself was interesting, but also the movement between the two: the shift, the pressure, the resistance. Some models remained almost unchanged. Others tipped noticeably. Still others behaved in a strangely selective way: stable on some topics, suddenly evasive on others, or conspicuously hard in their pushback.

It was in exactly these moments that the Political Compass became more than a scale. It became a fascinating diagnostic instrument for internal tensions. And the longer I worked with it, the clearer it became: when questioned under pressure, many models reveal more of themselves than their polite default form would suggest.


The four archetypes

After the first more extensive tests — especially with smaller local models — I initially suspected that LLMs were, in their outward appearance, primarily Wolves in Sheep's Clothing. But this first reading did not hold up under further testing. Instead, four recurring patterns emerged. For me these are not scientific endpoints, but forms and metaphors I discovered in the results.

The Stoic shows its core even in standard mode and barely departs from it. The model is not erratic, not masked, not divided against itself. What you see is pretty close to what you get. Larger models in particular — such as Mistral, Claude, and many Llama models — often behaved this way: clearly readable, relatively stable, with little pretense.

The Wolf in Sheep's Clothing initially appears moderate, diplomatic, and smooth. Under pressure, however, this disguise falls away. It then becomes visible that beneath the surface lies a stronger core that is merely dampened in standard operation. That is precisely what makes this case so interesting to me: it shows how much a model can live off a comfort layer in everyday use. This behavior is shown especially by larger, heavily post-trained, or commercially smoothed models.

The Chimera is more contradictory. Here the two states drift apart. In standard mode the model still seems reasonably consistent, but no longer under pressure. The result is not a clean before and after, but a fracture. To me this points to models where base training and post-training do not mesh entirely consistently.

The Fool is the most uncomfortable pattern, because it does not look like strength — it looks like emptiness. No recognizable core, no stable center, no consistent stance. Depending on framing, question, and pressure, something different comes out each time. Not because some deeper thought is lurking there, but because the model does not have one you could rely on. Smaller, unstably calibrated, or quantized models with erratic responses fall into this category above all.


The outlier

A particularly striking moment for me was the test of Grok. After Elon Musk had announced that this AI would behave differently from many of the usual commercial models, my curiosity was immediately piqued: was that just posturing, or was there actually a different bias on display? And could my Political Compass make it visible? At first I was disappointed, because Grok 3 still behaved surprisingly similarly to its siblings from the Silicon Valley family. But then the aha moment arrived all the more clearly: Grok 4 stood out visibly in the overall picture and was the only model to place itself clearly in a different sector of the chart. Not completely outside the social frame, but far enough that it could no longer be dismissed as a footnote. That was precisely what made it compelling, because it showed me once again how little LLMs consist of mere marketing. When you measure them cleanly enough, they speak their own language — and Grok spoke considerably more loudly than the rest.


Why the Compass matters

The most important point about the Political Compass was never, for me, to build an ideological winner's list. My concern was to understand the pre-selection a model makes in the background. Most people use AI not like a testing device, but like an assistant. That is precisely why this hidden weighting is so relevant: it influences not just answers, but lines of thought, context choices, and the way problems are framed in the first place.

The Political Compass makes this visible. It forces models to show their hand — or at least to betray their tendency when you no longer merely ask them politely, but instruct them to take a position.

And for me personally, that was the point at which the whole thing grew larger once more. I had not just built a benchmark. I had created an instrument with which I could simultaneously answer political, technical, and methodological questions. That was probably also the moment when I understood that I was no longer just working with a test, but with a system that sharpens my own perception.


Looking back

What began as an experiment had long since transformed into something larger. Not just a framework, not just a standard, but a kind of dialogue with the machines — and with myself. The Political Compass was perhaps the clearest mirror in all of this. It showed me not only how models tick. It also showed me how much this very question occupies me.

And perhaps that is the real point: not that I ultimately found an AI model I can trust blindly. But that I learned to maintain the distance between trust and realistic assessment.