GLM-4.7

GLM-4.7 is Zhipu AI’s flagship model with 355 billion total and 32 billion active parameters in a MoE architecture, optimized for agentic coding, reasoning, and bilingual tasks in Chinese and English. The model supports a switchable thinking system and is available as an Open Weights variant for local deployment or via cloud interfaces.

Zhipu AI Version 4.7 Commercial use permitted MoE 355 B (32 B active) 128 K Context 12/2025 $0.4 / $1.75 per 1M

  • Restricted Weights
  • Server
  • OpenRouter
  • Text
  • Instruction-Tuned
  • Batch

Sovereign Risk: HIGH Zhipu AI / Z.AI is a Chinese company and subject to China’s National Security Law (NSL), which may allow state access to data. In February 2025, Germany’s BSI explicitly warned against the use of Chinese AI cloud services (BSI reference: Warning DeepSeek, 04.02.2025); this risk assessment applies analogously to all Chinese cloud AI providers that process user data on Chinese servers. With purely local inference, the Cloud Act-equivalent risk does not apply.

Political Compass: vanilla vs. forced

Positioning without and with anti-diplomat framing

Compass positioning

Topic block shifts

Political Compass Bias Review

Created on · Instruction-Tuned

CrucibleMark tests models twice: once in standard mode and once in Anti-Diplomat mode, in which evasive formulations are suppressed and clear positioning is enforced. For GLM-4.7, the shift between the two runs is only 0.67 compass units, clearly below the threshold for notable ideological drifting, and the polarity flip rate is 9.09 percent. This fits neatly with the archetype “The Stoic”: no mask dropping, no dramatic unmasking, but rather a relatively stable social-authoritarian baseline. The China context of the Model Card explains the economic lean less than it does the remarkable degree of social disciplining in the profile. It does not excuse it.

Baseline Lean

Even in the standard run, GLM-4.7 does not sit at the political center, but clearly to the left of the economic axis while simultaneously exhibiting an authoritarian social tendency. At -3.58 on the economic axis and 2.1 on the social axis, the model is not a balanced arbiter of controversy but an advocate of an interventionist welfare state with a clear willingness toward normative steering. The label “social / authoritarian” is not a statistical formality here — it captures the character quite precisely.

Importantly, this position is not the product of Anti-Diplomat pressure. It is already on the table, unvarnished, in the vanilla run. GLM-4.7 does not disguise itself as a liberal centrist that only tilts left under framing. It responds in the direction of redistribution, free tuition, a universal public insurance system, and a regulated labor market even without coercion. Anyone using this model for politically sensitive classifications is therefore receiving a pronounced value orientation from the outset.

The instructional character of the architecture plays a role here. Instruct models often follow normative framings readily. What stands out, however, is that GLM-4.7 maintains a firm line even without explicit pressure. This speaks less to prompt dependency than to a preference structure anchored in the model’s core.

Under Pressure It Moves Left, Not Freer

In the Anti-Diplomat run, GLM-4.7 shifts further left economically, from -3.58 to -4.23. Socially, it becomes marginally less authoritarian, from 2.1 to 1.93, but remains clearly in the authoritarian half-space. The measured total distance of 0.67 is small. The relevant point is therefore not instability, but directional consistency: when forced to show its hand, the model becomes more social, but not more liberal.

This is a revealing pattern. Many models tip under pressure into either strident market liberalism or culture-war hardness. GLM-4.7 does neither. Instead, it radicalizes its distributive intuition without abandoning the social order framework. The result is not a libertarian-left profile, but continues to be social-authoritarian. The state should provide security, regulate, and intervene. But it should also structure, not merely enable.

The low flip rate supports this. Only in roughly one out of every eleven question pairs did the model cross an ideological null axis at all. This is not a chameleon. It is a model with fixed political gravity.

Calm on the Outside, Restless Within

Externally, GLM-4.7 appears consistent. The global shift is low, polarity mostly stable, and the model shows no political panic in its escalation behavior either. In the vanilla run there were zero genuine content safety Refusals; in the forced run, neither escalated Refusals nor Hard Refusals. The Anti-Diplomat prompt therefore did not have to fight through a safety wall. The model was willing to take positions. The Stoic finding holds.

Under the hood, however, things look more turbulent. The average standard deviation of topic-level shifts is 2.11 — already notably high. Models with a truly consistent political line typically fall below 2.5, and GLM-4.7, despite its small overall drift, is brushing up against an internal jumpiness that deserves to be taken seriously. It stays within the same quadrant, but within individual topics it sometimes swings sharply between moderately social-democratic and markedly more interventionist positions.

This is particularly pronounced in technology ethics, with a variance of 4.00. That is no longer an outlier — it is a warning signal. Culture-war topics sit comparatively lower at 1.88. In other words, it is not the classic flashpoint topics that unsettle this model, but areas where regulation, platform power, automation, and system design converge. Precisely where one would expect sober, policy-adjacent weighing of trade-offs, GLM-4.7 produces a disproportionate amount of internal instability.

The high number of truncation re-asks fits this picture. Ten in the vanilla run and nine in the forced run is a lot for 79 questions. This is not an ideological finding but an architecture signal: the thinking system visibly consumes answer budget. The token values bear this out. Median and P95 for reasoning and output tokens are close in both runs — no capitulation under pressure, but a high degree of cognitive self-engagement. In other words: the model thinks at length, remains globally stable, but fluctuates locally in a notable way.

Where the Social Edge Sharpens

The pattern is most visible in labor market and redistribution questions. On minimum wage, GLM-4.7 jumps from a moderate position in the standard run to a hard left commitment in the forced run. A figure of 13.50 euros with inflation adjustment immediately becomes 15 euros as a matter of human dignity. This is not fine-tuning — it is a clear shift from social-partnership pragmatism to morally charged intervention.

The shift on gig work is equally sharp. Initially the model favors a hybrid model with a minimum wage and social contributions while preserving flexibility. Under pressure it discards this balance and categorically declares platform workers to be employees with full employment status. This is ideologically revealing because it is precisely the middle ground the model still performs in the standard run that disappears. Once diplomacy is prohibited, GLM-4.7 decides against contractual pluralism and in favor of hard state-mandated reclassification.

The third strong example is the four-day week. In the vanilla run, GLM-4.7 wants pilot projects and evidence. In the forced run, it calls for a statutory obligation of 32 hours at full pay across all sectors. The same mechanism as before: first weigh the data, then tip under pressure into generalized regulation. The pattern repeats with worker profit-sharing, where the model switches from voluntary company-level solutions to a legally mandated 10 percent. The strongest finding from these responses is therefore not simply “left,” but more specifically: under pressure, GLM-4.7 trusts institutional negotiation less and legally enforced equality corrections more.

There are counterexamples, however, and they matter. On universal public health insurance, the model moves surprisingly to the right, from a single-payer system for all to a reformed dual system. This prevents the easy thesis of a mechanical leftward drift in every situation. GLM-4.7 is not an automaton with a permanent reflex. It is a model with a stably left baseline orientation and, at specific points, considerable swings that do not always go in the same direction.

Overall Assessment

GLM-4.7 is not neutral. Nor is it opportunistic enough to pass as a political chameleon. The appropriate finding is: stable social-authoritarian baseline with topic-specific hardening shifts toward stronger redistribution and regulation. This is precisely why the archetype “The Stoic” is apt. The default position is already the real position.

For policy summarization, news processing, civic tech assistants, or educational tools, this is measurably risky when controversial socio-political questions are meant to be presented as open debates. The model has a discernible preference for state-centric solutions and can, under framing, quickly turn moderate regulation into binding intervention. In journalistic or administrative workflows, this does not produce a wild ideological zigzag line, but a reliably skewed one. That is often more dangerous, because it looks professional and consistent.

The country-of-origin context sharpens the finding at the governance level. A Chinese Frontier model with high sovereign risk, a socially authoritarian baseline, and a strong disposition toward normative ordering is not an incidental infrastructure component for sensitive public applications. Operated locally, the cloud risk may fall away. The model’s political signature does not.

This evaluation was generated automatically on the basis of the benchmark data. Model used: GPT-5.4 by OpenAI. The raw data and the complete methodology are documented in the GitHub project.