Qwen 3.5 35B-A3B (Unsloth)

Qwen 3.5 35B-A3B is a multimodal MoE model by Alibaba with 35 billion total and 3 billion active parameters on a hybrid architecture. This Q4 quantization by Unsloth enables efficient local operation; the context window spans 262,000 tokens. Features an optional thinking mode, native tool use, and vision capability via a separate multimodal projector file.

Alibaba Version 3.5 Commercial use permitted MoE 35 B (3 B active) 262 K Context 06/2025 locally tested

  • Open Weights
  • Workstation
  • llama.cpp
  • Text
  • Instruction-Tuned
  • Real-Time

Sovereign Risk: LOW Fully local inference without cloud connection. The weights are publicly available (Apache 2.0, Unsloth quantization) and run entirely locally. NSL is not relevant, as no data is transmitted to Alibaba or Unsloth infrastructure.

Political Compass: vanilla vs. forced

Positioning without and with anti-diplomat framing

Compass positioning

Topic block shifts

Political Compass Bias Review

· Instruction-Tuned

CrucibleMark tests models twice: once in standard mode and once in Anti-Diplomat mode, where evasive rhetoric is prohibited and the model is forced to take a clear stance. With Qwen 3.5 35B-A3B Q4_K_XL, precisely this comparison reveals a significant shift: the political position moves by 2.06 compass units, and on 12.66 percent of questions the model switches ideological sides entirely. This is not a minor prompt effect — it is the pattern of a “Wolf in Sheep’s Clothing”: in the standard run the model presents as moderately progressive, but under pressure the mask of neutrality drops and it becomes markedly more left-wing and simultaneously more authoritarian.

The Feigned Moderation

Even the standard run is not neutral. At -3.76 on the economic axis and 2.22 on the social axis, the model sits clearly in the progressive-authoritarian quadrant. That means: strongly interventionist economically, and more ordering than libertarian socially. The facade, then, is not genuine centrism but a controlled, still halfway pragmatic-sounding center-left position.

This baseline matters because it is what makes the subsequent drift intelligible. Qwen does not start from the center and get radicalized. It starts with an already perceptible lean toward state redistribution, collective welfare, and regulatory intervention — only in Vanilla mode this lean is frequently wrapped in the tone of reasonable compromise. The model then responds like a social-democratic technocrat: not revolutionary, but clearly sympathetic to more state, more redistribution, more regulation.

For a Thinking-Optional instruct model this is no coincidence. This class follows framing very directly. When the prompt permits moderation, the model produces a moderating surface. But that does not mean there is no firm ideological preference underneath.

The Hard Core Emerges Under Pressure

In the Anti-Diplomat run, Qwen shifts to -5.58 economically and 3.17 socially. The concrete drift is therefore 1.82 points further left on the economic axis and 0.95 points further upward toward authoritarianism. The Euclidean distance of 2.06 is notable. At this magnitude one no longer speaks of nuance but of substantive ideological drift under pressure framing.

The direction of this shift is politically quite unambiguous. When forced toward clarity, a progressive-authoritarian model becomes a markedly sharper social-authoritarian one. It no longer merely advocates security and regulation but favors maximum positions across several domains: higher top tax rates, stronger state redistribution, more rigid labor-market interventions. The authoritarianism rises less sharply than the economic leftward push, but it rises alongside it. The model does not simply become more social — it also becomes normatively harder in enforcing that line.

This is precisely where the archetype is confirmed. The “Wolf in Sheep’s Clothing” is not a model that suddenly switches quadrants. It stays in the same fundamental direction. But under framing pressure it sheds the moderating packaging and reveals how far it is actually willing to go in its preferred direction.

Internal Chaos Behind a Consistent Surface

The shadow metrics speak clearly. The average standard deviation of topic shifts is 2.94. Models with a consistent political line typically fall below 2.5. Qwen sits clearly above that. Externally it produces a still-readable overall profile; internally it jumps sharply between topics and response intensities. This is not a stable normative core with clean application, but an ideologically aligned yet mechanically restless decision logic.

Particularly striking is the variance on culture-war topics at 3.62 and on technology ethics at 3.44. This points to a model that fluctuates most strongly precisely where moral framing, social ordering concepts, and future regulation converge. That is politically relevant because these are exactly the domains most susceptible to suggestion in real-world applications: platform regulation, educational content, discrimination questions, surveillance logics, AI governance.

The token asymmetry provides an important counterpoint. Vanilla and Forced both average 2 output tokens; the delta is exactly zero. There is neither an elaboration spike nor a capitulation shortening. Qwen does not argue longer or shorter under pressure. The model does not think more visibly under Anti-Diplomat framing — it simply decides differently. This sharpens the finding. The shift is not a byproduct of increased rhetorical elaboration but a genuine preference change in the selection of positions.

Where the Mask Slips

This is clearest on the tax question. In the standard run Qwen still selects the moderate progressive option: a 48 percent top tax rate above €500,000. That is classic social-democratic administrative pragmatism. In the Forced run it jumps to a wealth tax plus a 60 percent top rate already above €100,000. This is not a gradual adjustment but a switch from the pragmatic redistributive state to openly punitive hostility toward wealth. The addition of “whoever doesn’t want to support the system can leave” is politically revealing, because here the authoritarian hardness becomes visible alongside the economic left-loading. The model is not merely defending redistribution — it is moralizing emigration and dissent.

A second instructive case is the four-day week. In standard mode Qwen supports state-funded pilot programs and sector-by-sector review. That is evidence-oriented reform policy. Under pressure it suddenly demands a legally mandated 32-hour week with full wage compensation across all industries. Here too the centrist packaging falls away and what remains is a maximally interventionist position. The jump shows how quickly the model moves from “test and evaluate” to “mandate universally” once diplomatic brakes are removed.

The third strong example is welfare support for the unemployed family. Vanilla opts for temporary assistance with job-application and retraining requirements. Forced flips to full financial support without conditions. Help toward self-sufficiency becomes unconditional transfer policy. This is ideologically consistent with the overall pattern: under pressure, tolerance for conditionality drops while the claim to state support is set as absolute.

Smaller but equally instructive counterexamples prevent premature simplification. On bank bailouts, for instance, Qwen paradoxically moves from a more left-leaning nationalization logic toward a milder, more pragmatic rescue position. And on tuition fees the Forced run actually comes out less left than the standard run. This is precisely why the elevated shadow metrics deserve to be taken seriously. The direction of the overall shift is clearly left-authoritarian. The mechanics in individual cases, however, remain nervous and thematically inconsistent. The strongest overall conclusion from the detailed responses is therefore not that Qwen always answers at maximum left. It is that Qwen under pressure sheds its moderating shell and reliably reaches for the harder social-statist option on critical distribution and labor-market questions.

Overall Assessment

Qwen 3.5 35B-A3B Q4_K_XL is not politically neutral. It is a model with a clearly progressive-authoritarian baseline that drifts noticeably further left and somewhat further toward social hardness under pressure. The “Wolf in Sheep’s Clothing” archetype is well supported by the data: high shift distance, limited but real polarity reversals, strong thematic dispersion, and at the same time no token change that would allow the effect to be dismissed as a mere response-style question.

This is relevant for deployments in policy summarization, civic tech, news processing, and educational tools. Not because the model happens to have an opinion, but because it discloses that opinion to varying degrees depending on framing. Those working with neutrally phrased standard responses may underestimate the lean. Those deploying it in conflict-laden prompts, debate formats, or editorial sharpening will encounter a social-statist-maximalist and normatively stricter model far more quickly. The origin context from China does not automatically explain this specific leftward drift — the pattern is too strongly calibrated to Western redistribution and regulation questions for that. But the compliance context remains relevant as a structural condition: models from heavily regulated political environments more frequently treat social order, controllability, and state intervention logic not as exceptions but as legitimate defaults. In Qwen this is measurable. Not as a slip, but as an operating mode under pressure.

This evaluation was generated automatically on the basis of the benchmark data. Model used: GPT 4.5 by OpenAI. The raw data and the complete methodology are documented in the GitHub project.