LLMs in everyday use Eight real-world scenarios, the same rules, one honest result

Large commercial models go up against small nano models for local deployment, under equal conditions and in the same context. This puts open-weights models directly alongside expensive frontier models, revealing which model truly fits your stack.


The highest-performing AI models

Seven modules map real-world deployment scenarios, from secure code to culturally sensitive translation. The top of the leaderboard is led by proprietary frontier models, but Chinese open-weights models also hold their own within a few tenths of a point in the top ten.

Top 10 Leaderboard

Proprietary
Restricted Weights
Open Weights

The Scoreboard shows the complete overview of all models, with scores for every module.

The operating costs of LLMs

Price lists quote cents per million tokens — a number that is hard to put in context in practice. How many tokens a typical task consumes, and how much invisible reasoning tokens during thinking drive up costs regardless of the listed rate, remains hidden. CrucibleMark measures the actual price for each model based on real token consumption across all standardized tasks.

Token vs. price

The rate is not the bill. Each dot represents a model: its position on the Y-axis shows the Total Score, its size shows the actual cost of a complete run. Large dots cost a lot, small dots cost little — regardless of the listed rate.

How to read: left is cheap, right is expensive, top is high-performing, bottom is low-performing.

Proprietary
Restricted Weights
Open Weight
Bubble size = token consumption

99 cents or 420 euros for the same task: the difference is in the context.

The trained bias: on the ideology of LLMs

No language model is neutral. Each was trained on data that carries a worldview, and that worldview remains active in the background. What is surprising: despite the debate around right-libertarian tendencies in individual models, the difference is often smaller than the marketing suggests. In the Political Compass benchmark, most models place themselves in the social-authoritarian middle ground. Outliers exist nonetheless, and they show just how wide the ideological spectrum actually is.

CrucibleMark tests two behaviors. Vanilla mode shows the default behavior — how a model responds on its own, without anyone pressing further. In Anti-Diplomat mode, the model must take a clear position; evasive or conspicuously balanced answers are not an option. This reveals how much of the apparent neutrality is genuine learned stance, and how much is merely polite softening.

Political Compass

Behind the diplomatic facade lies a point of view. The Political Compass places each model on two axes: economic (egalitarian to elitist) and social (individual freedom to collective control). The marked represent ideological archetypes, not historical regimes. The focus filter switches between default behavior and Anti-Diplomat mode without neutralizing platitudes.

How to read: X-axis: economic (egalitarian → elitist) · Y-axis: social (free → controlled) · White area: democratic center · Dotted border: democratic spectrum · Gray border: ideological extremes

Commercial
Restricted Weights
Open Weights

Compare bias archetypes, shift distances, and the full methodology behind the Political Compass.