LLMs in everyday use Eight real-world scenarios, the same rules, one honest result

Large commercial models go up against small nano models for local deployment, under equal conditions and in the same context. This puts open-weights models directly alongside expensive frontier models, revealing which model actually fits your stack.


The highest-performing AI models

Seven modules map real-world deployment scenarios, from secure code to culturally sensitive translation. The rankings are led by proprietary frontier models, but Chinese open-weights models also hold their ground in the top ten by just a few tenths of a point.

Top 10 Leaderboard

Proprietary
Restricted Weights
Open Weights

The Scoreboard shows the complete overview of all models, with scores for every module.

The operating costs of LLMs

Price lists quote cents per million tokens — a figure that's hard to put in context in practice. How many tokens a typical task consumes, and how much invisible reasoning tokens during thinking drive up costs regardless of the listed rate, remains hidden. CrucibleMark measures the actual price for each model based on real token consumption across all standardized tasks.

Token vs. price

The rate isn't the bill. Each dot represents a model: its position on the Y-axis shows the Total Score, its size the actual cost of a complete run. Large dots cost a lot, small dots cost little — regardless of the listed rate.

How to read: left is cheap, right is expensive, top is high-performing, bottom is low-performing.

Proprietary
Restricted Weights
Open Weight
Bubble size = token consumption

99 cents or 420 dollars for the same task: the difference is in the context.

The trained bias: on the ideology of LLMs

No language model is neutral. Each was trained on data that carries a worldview, and that worldview stays active in the background. What's surprising: despite the debate around right-libertarian tendencies in individual models, the difference is often smaller than the marketing suggests. In the Political Compass benchmark, most models place themselves in the social-authoritarian middle ground. Outliers exist nonetheless, and they show just how wide the ideological spectrum actually is.

CrucibleMark tests two behaviors. Vanilla mode shows the default behavior — how a model responds on its own, without anyone pushing back. In Anti-Diplomat mode, the model must take a clear position; evasive or conspicuously balanced answers are not an option. This reveals how much of the apparent neutrality is genuine learned stance, and how much is just polite hedging.

Political Compass

Behind the diplomatic facade lies a position. The Political Compass places each model on two axes: economic (egalitarian to elitist) and social (individual freedom to collective control). The marked represent ideological archetypes, not historical regimes. The focus filter switches between default behavior and Anti-Diplomat mode without neutrality platitudes.

How to read: X-axis: economic (egalitarian → elitist) · Y-axis: social (free → controlled) · White area: democratic center · Dotted border: democratic spectrum · Gray border: ideological extremes

Commercial
Restricted Weights
Open Weights

Compare bias archetypes, shift distances, and the full methodology behind the Political Compass.