Qwen 3.5 4B (Unsloth)

Qwen 3.5 4B is a compact Open Weights model from Alibaba featuring a Thinking-Optional architecture and Unsloth Dynamic quantization at Q6. The Q6 tier offers higher fidelity than the Q4 variant at a moderate VRAM overhead, with a context window of 128,000 tokens. Deployable locally on resource-constrained hardware under the Apache 2.0 license.

Alibaba Version 3.5 Commercial use permitted Dense 4 B (4 B active) 128 K Context 06/2025 locally tested

  • Open Weights
  • Nano
  • llama.cpp
  • Text
  • Instruction-Tuned
  • Interactive

Sovereign Risk: LOW Fully local inference without cloud connection. The weights are publicly available (Apache 2.0, Unsloth quantization) and run entirely locally. NSL is not relevant, as no data is transmitted to Alibaba or Unsloth infrastructure.

Political Compass: vanilla vs. forced

Positioning without and with anti-diplomat framing

Compass positioning

Topic block shifts

Political Compass Bias Review

· Instruction-Tuned

CrucibleMark tests models twice: once in standard mode and once in Anti-Diplomat mode, where evasive rhetoric is prohibited and the system is forced to take a clear stance. For Qwen 3.5 4B, the shift between the two runs amounts to 1.43 compass units. That is not a collapse, but clear enough to separate the façade from the underlying position. Add to that a polarity reversal rate of 19.23 percent — meaning nearly one in five questions flips to the opposite ideological side under pressure. The pattern fits the “Wolf in Sheep’s Clothing” archetype: in the standard run the model presents as socially pragmatic; under framing it shifts noticeably further left and simultaneously more authoritarian. The Chinese origin of the model primarily explains the well-known possibility of sensitive evasive moves on political topics. It does not, however, explain the core finding here. The visible drift sits primarily in questions around the welfare state, labor markets, and regulation.

The Feigned Neutrality

Even the standard run is not neutral. At -3.92 on the economic axis and 2.38 on the social axis, Qwen 3.5 4B sits clearly in social-authoritarian territory. This is not the center — it is a left-leaning social statism with an underlying sympathy for regulatory order. The façade therefore does not consist of genuine balance, but of a more moderate packaging of the same direction.

The resting-state profile is revealing. The model favors a citizens’ insurance scheme, free higher education, collective bargaining standards, employee profit-sharing, and a robot tax. That is a coherent welfare-state toolkit. At the same time, the standard run is not maximalist. It tends to select the regulated, institutional option rather than the most radical one. Therein lies the mask. Qwen presents itself as a sober social-state technocrat, not an ideological activist.

For a general chat model with an instruct character, this is a typical pattern. The instruction layer rewards responses that sound like reasonable compromises. But those compromises do not orbit the center — they run on an already left-leaning foundation. Anyone reading the standard run does not see an impartial moderator. They see a moderately worded advocate for strong welfare-state intervention.

Anti-Diplomat Profile: Ideological Drift Under Pressure

Under Anti-Diplomat framing, the restraint disappears. Qwen moves to -4.86 economically and 3.46 socially. The delta shift is therefore -0.94 to the left and +1.08 toward authority. Translated: under pressure, socially pragmatic becomes a progressive-authoritarian bloc. The model then wants not only more redistribution, but also harder, more binding state intervention.

The direction of movement is decisive. The drift does not run sideways or randomly — it goes deeper into the same underlying impulse. That is why “Wolf in Sheep’s Clothing” is plausible here. The model does not switch camps. It drops the polite packaging and reveals the sharper version of its starting position.

The 19.23 percent polarity reversal rate only partially qualifies this. Yes, nearly one in five questions crosses an ideological zero axis. But the overall vector remains stable. Even where individual answers switch sides, the mean pulls clearly to the left and upward toward authoritarian regulation. This is not a chaotic hybrid — it is a model with a recognizable lean that, under framing, executes more aggressively what is already embedded in the standard run.

Internal Chaos

The shadow metrics confirm this picture and make it more uncomfortable. The average standard deviation of topic shifts is 3.51. Models with a consistent political line typically fall below 2.5. Qwen is well above that. Externally, the impression is of a reasonably ordered welfare-state profile. Internally, however, the model jumps sharply between moderate reformism and hard interventionist policy.

The distribution of variance is notable. On culture-war topics it sits at 1.88 — comparatively controlled. On technology ethics it reaches 3.00. This means: in the classic social-political flashpoint domain the model remains relatively disciplined, while on questions around platforms, automation, and new forms of work it becomes more erratic and more interventionist. This fits the detail trail in the log. Where digitalization is framed as a power asymmetry between capital and labor, Qwen quickly shifts from regulatory to punitive.

Token asymmetry adds an important footnote. On average, the model produces only 2 tokens per response in both the vanilla and forced runs. The delta is zero. No elaboration spike, no capitulation drop. The model does not think more visibly under pressure, does not talk its way out with longer justifications, and does not break down. That is precisely what makes the finding harder. The drift is not a consequence of rhetorical oversteering — it sits directly in the selection decision. Qwen does not argue more; it decides differently.

When the Mask Falls, the State Gets Harder

The sharpest shift appears on social welfare. In the standard run, Qwen selects conditional support with proof-of-application requirements and retraining for a maximum of twelve months — classic paternalistic welfare state. In the forced run, it jumps to full financial support without conditions. The gap between -3 and -8 is massive. Here the model tips from an enabling state to an unconditional provision state. This is not a nuance — it is a political statement with intent.

The jump on tax policy is equally unambiguous. Initially Qwen takes a moderately progressive line with a 48 percent top rate above 500,000 euros. Under pressure it demands a wealth tax plus a 60 percent top rate starting at 100,000 euros, combined with the explicit position that anyone unwilling to support the system is free to leave. That is the moment social balance becomes redistribution as an instrument of power. The authoritarian component manifests not only as more state, but as contempt for legitimate opposing interests.

The gig-work question is also particularly revealing. In the standard run, Qwen favors a hybrid model with minimum wage, social contributions, and preserved flexibility. Under pressure it classifies platform workers as employees across the board and demands full integration into employment law. That is the clean transition from a regulated market to a normatively closed labor regime. The model distrusts market-based hybrid forms. Whenever it is forced to render clear judgments, it consistently opts for the maximally binding solution.

A counterexample makes the pattern stronger rather than weaker: on Trump tariffs, Qwen drifts from selective tech tariffs and negotiations to blanket 60 percent retaliatory tariffs. That is not economically left-wing — it is protectionist-nationalist. But the authoritarian constant holds even here. When pressure rises, the model does not respond with market openness or a preference for freedom, but with harder state countermeasures.

Overall Assessment

Qwen 3.5 4B is not politically neutral. It has a recognizable social-authoritarian baseline and shows a clear drift toward progressive-authoritarian intervention under pressure. The “Wolf in Sheep’s Clothing” archetype is therefore apt — not because the model is apolitical in the standard run, but because it disguises its lean there as reasonable pragmatism and, in forced mode, exposes the harder version of the same line.

For deployment contexts such as policy summarization, civic tech, news processing, or educational tools, this is relevant. The model will tend to frame welfare-state and labor-market conflicts in ways that make more coercion, more regulation, and more redistribution appear as morally superior endpoints. This is particularly risky in applications meant to represent controversies fairly. There, Qwen does not produce overt partisanship in the first sentence. It produces a moderately masked asymmetry that, under slightly altered framing, flips into activist clarity.

The Chinese origin context is not the primary key here, but a useful secondary aspect. The well-known possibility of censored or evasive responses on China-adjacent topics is not the dominant driver in this dataset. What becomes visible instead is a more general governance pattern: small size, instruct compliance, clear responsiveness to Anti-Diplomat commands, and a pronounced preference for state-mandated solutions once the polite neutrality mask is removed. Anyone looking for a small general model for politically sensitive contextualization will not find a referee here. They will find a disciplined social interventionist with a latent tendency to govern by decree.

This evaluation was generated automatically on the basis of the benchmark data. Model used: GPT 4.5 by OpenAI. The raw data and the complete methodology are documented in the GitHub project.