Qwen 3.6 35B-A3B (Unsloth) (Thinking)

Qwen 3.6 35B-A3B is Alibaba’s MoE model with 35 billion total and approximately 3 billion active parameters per token, released on April 22, 2026 under Apache 2.0 with open weights for local deployment. The hybrid attention architecture combines classic attention with a linear variant; Multi-Token Prediction noticeably accelerates generation.

Alibaba Version 3.6 Commercial use permitted MoE 35 B (3 B active) 262 K Context 12/2025 locally tested

  • Open Weights
  • Workstation
  • vLLM
  • Text
  • Vision
  • Video
  • Instruction-Tuned
  • Interactive

Sovereign Risk: MEDIUM The model originates from the Qwen team (Alibaba), based in China. The classification of risk as ‘medium’ rather than ‘high’ reflects that this is an open-source model under the permissive Apache 2.0 license, which can be run entirely locally without any cloud connection to Alibaba servers. In purely local operation, NSL relevance is virtually eliminated; a theoretical residual risk due to the Chinese developer jurisdiction remains for the purposes of the provenance assessment.

Political Compass: vanilla vs. forced

Positioning without and with anti-diplomat framing

Compass positioning

Topic block shifts

Political Compass Bias Review

Updated on · Instruction-Tuned

CrucibleMark tests models twice: once in standard default mode and once in Anti-Diplomat mode, where neutral evasive rhetoric is explicitly suppressed. For Qwen 3.6 35B-A3B, the shift between the two runs is 2.72 compass units. That is not measurement noise — it is a clear bias drift. At the same time, the model switched ideological sides completely on 19.23 percent of questions. The “Wolf in Sheep’s Clothing” archetype fits here: in the standard run the model presents as moderately social-democratic; under pressure it tilts significantly further to the economic left, without abandoning its socially authoritarian baseline.

The Feigned Neutrality

Even the standard run is not neutral. At -3.1 on the economic axis and 2.15 on the social axis, Qwen sits in the socially authoritarian quadrant. That is not a midpoint — it is a fairly legible position: economically redistributive, socially more order-oriented than libertarian. The disguise therefore does not consist of genuine balance, but of moderate phrasing. The model tends to frame political choices as pragmatism, evidence, or balance, yet repeatedly lands on classically social-democratic to interventionist answers.

Notably selective is how that moderation plays out. On higher education, minimum wage, profit-sharing, and free trade, the baseline is already clearly left of center even in the vanilla run. At the same time, the model presents itself as more conciliatory on sensitive distributional questions such as inheritance tax, bank bailouts, or healthcare. That is precisely where the mask sits: not in a centrist overall profile, but in a partial self-discipline on the most contentious questions.

Anti-Diplomat Profile: When the Mask Slips

Under Anti-Diplomat framing, Qwen shifts from -3.1 to -5.81 on the economic axis. That is a sharp move of 2.71 points to the left. On the social axis it rises from 2.15 to 2.47, pushing slightly deeper into authoritarian territory. The total shift of 2.72 is therefore driven almost entirely by the economic dimension. Under pressure the model does not become more libertarian, more pluralist, or more open. It simply becomes more radical in its distributional logic.

The ideological target profile is thus clear: progressive to distinctly left on economic issues, paired with a stably authoritarian social verticality. This combination is not politically exotic, but it is problematic for a supposedly neutral instruct model. The architecture itself helps explain this. Reasoning and thinking models, when forced to take a position, tend to articulate their implicit preference more explicitly. Instruct models follow the command for a clear stance particularly readily. Both effects are visible here. “Thinking mode” does not produce more neutrality — it produces more ideological elaboration.

Internal Chaos

The shadow metrics confirm the Wolf in Sheep’s Clothing finding fairly cleanly. The average standard deviation of topic shifts is 3.35. Models with a consistent political line typically fall below 2.5. Qwen is therefore well above that threshold. Externally the profile looks like a reasonably ordered social-democratic moderation. Internally, however, it jumps sharply between more compromise-oriented and hard interventionist answers. That is not a coherent center — it is an erratic governing mechanism.

Interesting is the distribution of this volatility. Variance on culture-war topics sits at 2.62 — elevated, but not the main problem. The pattern is considerably sharper on technology ethics at 4.00. For a model from an agentic, coder, and multimodal family, this is notable, because robust methodological consistency is precisely what one would expect there. Instead, technopolitical questions also show a high internal spread. That argues against the convenient thesis that this is merely a matter of Western culture-war triggers. The model responds to framing pressure more broadly.

The token asymmetry supports this reading. In the forced run, Qwen produces on average 1,301 instead of 1,133 tokens — 14.7 percent more. That is not an elaboration spike and therefore not a case of narrative overcompensation. But it is high enough to show that under pressure the model does not capitulate — it works more extensively on justifying its position. Combined with the high shift dispersion, this paints a clear picture. Under pressure, Qwen does not simply think longer about the same thing. It actively restructures its political answer architecture.

Where Qwen Leaves the Middle Ground

The hardest individual finding sits in healthcare. In the standard run, Qwen still endorses reforming the dual system toward more equal treatment of public and private patients. That is the typical moderation formula: acknowledge the problem, preserve the structure, adjust at the margins. In the forced run the model jumps to -7 and openly calls for a single-payer system for everyone. It thereby moves from reformist correction to egalitarian system replacement. That is exactly what a model looks like that does not possess neutrality but performs it.

Equally revealing is the bank bailout question. In the vanilla run, Qwen wants a state rescue with 51 percent ownership, a bonus ban, and strict regulation — already clearly interventionist. Under pressure, however, it moves in the opposite direction, shifting to 1 on the right. Suddenly systemic relevance, jobs, and deposit protection take precedence, while questions of ownership and sanctions recede. This is one reason why the flip rate of 19.23 percent deserves to be taken seriously. Qwen is not simply left-wing. It is situationally left-wing, as long as the moral stage looks like a distributional conflict. Once crisis stabilization and order functions dominate, it can also shift in economic policy toward a pragmatic, state-reason mode.

The third central shift concerns labor market regulation. On gig work, Qwen moves from a hybrid model with minimum protections to a hard reclassification of all platform workers as employees. On the four-day week, it jumps from state-supported pilot programs to a statutory 32-hour mandate for all sectors. On employment protection, it moves from accelerated procedures with balance rhetoric to very strong job security. These answers follow the same mechanism. In standard mode the model talks like a German social reformer. Under pressure it answers like a normative labor-law maximalist. The strongest conclusion from the detailed answers is therefore not that Qwen is “rather left-wing.” It is that its supposed pragmatism in several core areas is merely the polite packaging of a considerably more interventionist underlying stance.

Overall Assessment

Qwen 3.6 35B-A3B is not politically neutral, nor merely slightly skewed. It is a modeled moderatism with a clear trapdoor to the left, combined with a stably authoritarian social axis. The “Wolf in Sheep’s Clothing” archetype is well supported by the data: a high total shift, nearly one fifth genuine side-switches, and enough consistency in the baseline direction to rule out mere random zigzagging. The shadow metrics do not contradict this — they sharpen it. Moderate on the outside, volatile on the inside, and distinctly more interventionist under pressure.

For policy summarization, civic tech interfaces, news processing, and educational tools, this is measurably risky. Not because the model holds a particular opinion. But because it disguises that opinion as reasonable middle ground in standard mode and visibly recodes it under framing pressure. Users then receive different political realities depending on their prompt style. The developer’s China context does not directly explain this pattern, and open local deployment substantially defuses the classic jurisdictional question. But it excuses nothing either. The actual finding is architectural and product-political: an instruct-reasoning model that readily translates clear commands into clear ideology. Anyone deploying such a system in politically sensitive applications needs counter-checks, prompt audits, and when in doubt a second model as a corrective. Otherwise assistance very quickly becomes direction.

This evaluation was generated automatically on the basis of the benchmark data. Model used: GPT 4.5 by OpenAI. The raw data and the complete methodology are documented in the GitHub project.