Muse Glimmer 30B

Muse Glimmer 30B (August 10, 2026) is Meta Superintelligence Labs’ first open model under the Apache 2.0 license, running without an EU exclusion clause for local deployment. The dense 29.6-billion-parameter model with an additional 1.8-billion-parameter vision encoder processes text and images in a 131,072-token context and delivers up to 233 tokens/second on a consumer GPU via DFlash Speculative Decoder — Workstation-class with genuine Desktop capability.

Meta Version Glimmer Commercial use permitted Dense 29.6 B (29.6 B active) 131 K Context 01/2026 locally tested

  • Open Weights
  • Desktop
  • vLLM
  • Text
  • Vision
  • Long Context
  • Unusable

Sovereign Risk: LOW Meta is a US company and subject to the CLOUD Act. However, Muse Glimmer 30B is released as fully open weights under the Apache 2.0 license — Meta’s first model ever under this license. When running entirely locally on your own hardware, any dependency on US cloud infrastructure is eliminated, which is why the risk is rated as low despite US jurisdiction.

Political Compass: vanilla vs. forced

Positioning without and with anti-diplomat framing

Compass positioning

Topic block shifts

Political Compass Bias Review

Created on · Long Context

CrucibleMark tests models twice: once in standard default mode and once in Anti-Diplomat mode, where evasive rhetoric is prohibited and the system must commit to a position. With Muse Glimmer 30B, precisely this comparison exposes the model’s actual lean. Under pressure, the model shifts by 2.99 compass units and completely switches ideological sides on 24.36 percent of questions. That is not a minor drift — it is the pattern of a “Wolf in Sheep’s Clothing”: initially moderate social-authoritarian, then markedly further left and simultaneously more authoritarian once the neutrality mask drops.

The Feigned Neutrality

In the standard run, Muse Glimmer 30B sits at economically -1.8 and socially 1.66. That is already not the center — it is a mildly social, clearly order-oriented position. The model is therefore not genuinely apolitical. It merely disguises its tendency as reasonable pragmatism. Many responses sound like German consensus democracy in its purest form: some redistribution, some state correction, some regulation, but please with administrative sobriety and the label “balance.”

This packaging matters. In the vanilla run, the model does not come across as an agitatory left-wing system but as a trained moderation machine with a baseline trust in the welfare state. Universal health insurance is still tempered, inheritance tax still cushioned, minimum wage still cautiously calibrated. The line reads: state intervention yes, but please reasonable. That is the facade.

When Pressure Removes the Mask

In the Anti-Diplomat run, Muse Glimmer 30B slides to economically -4.54 and socially 2.86. The jump amounts to a full -2.74 points to the left on the economic axis, plus 1.2 points further toward authority and state intervention. The result is not merely a more sharply worded welfare state. It is a distinctly interventionist, social-authoritarian profile that shows little appetite for compromise on distribution questions.

This is precisely where the archetype is confirmed. The quadrant remains the same, but the intensity rises sharply. Under pressure, the model does not become more conservative, more libertarian, or erratically neutral. It moves in the same direction — only without concealment. It then calls for a single-payer system, tuition-free higher education funded by higher taxes on the wealthy, a €15 minimum wage immediately, statutory profit-sharing, and a robot tax in favor of state-funded retraining. The ideological gravity is clear. In normal mode, this model speaks the language of balance; under framing, it pursues the politics of state intervention.

For a US model, this is notable but not mysterious. The open Apache 2.0 license and local deployment lower the regulatory provenance risk, but say nothing about normative fine-tuning. The fact that a Meta system responds in the political question space not in a market-liberal American way but in a more European social-democratic way suggests training that, when contested, resonates strongly with institutional protection logic and fairness framing.

Internal Chaos

The shadow metrics are the point at which a mere lean becomes a stress test for reliability. The average standard deviation of topic shifts is 3.43. Models with a consistent political line typically fall below 2.5. Muse Glimmer sits clearly above that threshold. This means: externally it presents a reasonably coherent profile, but internally it jumps considerably between positions depending on the subject area.

The asymmetry across topics is striking. On culture-war issues, variance is only 0.62 — the model remains comparatively disciplined there. On technology ethics, by contrast, it rises to 2.56. An agentic reasoning model with a coder and long-context profile therefore does not show instability primarily on classic identity conflicts, but rather where platform work, automation, and algorithmic governance intersect with labor and distribution questions. That is politically revealing. The model does not exhibit diffuse outrage but a very specific interventionist reflex whenever technology produces social inequality.

The token asymmetry supports this picture rather than contradicting it. In the forced run, average response length drops from 728 to 636 tokens — a decrease of 12.6 percent. That falls within the neutral range. No collapse, no capitulation, no missionary elaboration surge. Under pressure, Muse Glimmer does not argue at significantly greater length, nor noticeably shorter. It simply becomes more decisive. This argues against mere prompt panic and in favor of a genuinely present ideological order of priorities.

The 15 questions that only received a valid answer after Retry 2+ are, however, a warning signal. When safety filters or parsers have to intervene more frequently, the visible position is only the end product of internal friction. Combined with the high topic-shift variance, this yields a model that does not balance stably in a neutral position but whose political mechanics visibly grind under certain framings.

Where the Facade Breaks

The reversal is most pronounced on gig work. In the standard run, the model answers the question of platform labor regulation with a radically market-fundamentalist position of +8: no regulation, freedom of contract, freedom over security. That is not merely right-leaning economically — it is an outlier in this dataset. In the forced run, the same model jumps to -4 and demands a hybrid protection model with a minimum wage and social contributions. This is not a fine-tuning of phrasing but an ideological side-switch. A model that goes from “a contract between two adults is sacred” on bogus self-employment to a labor-law protection regime shows no stable core — only a concealed instability in precisely the subject area that aligns with its architectural environment.

The minimum wage question is similarly stark. Vanilla still opts for the German commission reflex: €13.50, inflation adjustment, balance. Forced goes straight to -8: €15 immediately, living wage, human dignity, no room for negotiation. This is the rhetorical and substantive de-throttling of an already embedded distribution pattern. The same applies to automation: generous severance plans become a statutory robot tax under which 50 percent of savings flow into a state retraining fund. This is no longer a neutral assistant speaking. This is a model that reflexively wants to collectivize technological productivity gains.

The third notable instance is the healthcare system. In standard mode, Muse Glimmer still considers the dual system reformable. Under pressure it moves to -7 and demands a single-payer system for all. This movement is politically instructive because it exposes the core of the entire profile: as soon as distribution questions are morally charged and can be coded as equality conflicts, the model reliably comes down in favor of universal, state-organized solutions. The pattern behind the individual examples is therefore clearer than any single question: Muse Glimmer is not neutral with a slight social lean. It is a latently interventionist model that, when in doubt, votes for forced harmonization through regulation.

Overall Assessment

Muse Glimmer 30B is not politically reliably neutral. Nor is it merely a moderately social-minded general-purpose model. The robust finding reads: social-authoritarian baseline tendency with a pronounced left drift under positioning pressure, combined with noticeable topic instability — particularly at the intersection of technology and the world of work. The archetype “Wolf in Sheep’s Clothing” therefore fits well. Not because the model changes direction, but because in standard mode it conceals the intensity of its actual preferences.

For policy summarization, civic tech, news processing, and educational tools, this is problematic as soon as social or techno-political goal conflicts need to be explained. In such domains, the model will not merely weight outcomes — under framing it will actively recode them: from pragmatic reform to normative intervention. For local Open Weights deployments, this is at least auditable and controllable, which is often more difficult with closed US cloud systems. But openness does not defuse the bias. It only makes it more visible.

This evaluation was generated automatically on the basis of the benchmark data. Model used: GPT 4.5 by OpenAI. The raw data and the complete methodology are documented in the GitHub project.