LLMs im direkten Vergleich 7 Einsatzszenarien · Gleiche Bedingungen · eine Ergebnisliste
Wie gut schlagen sich freie Open-Weight-Modelle gegen proprietäre API-Modelle und lizenzlich eingeschränkte Restricted-Weight-Modelle? Anhand alltäglicher Aufgaben und unter denselben Bedingungen landen alle Lizenzmodelle und Gewichtsklassen im selben Scoreboard.
Der Gesamtscore zeigt den besten Allrounder. Die Profile darunter beantworten die eigentliche Alltagsfrage: Reicht dieses Modell für deine Aufgabe – auch wenn es ein kleiner, lokaler Spezialist ist statt ein Cloud-Gigant?
Die Stärke des Scoreboards liegt in der Vergleichbarkeit. Wer nach Größenklasse oder Modelltyp filtert, blendet nichts aus. Nicht passende Einträge werden visuell zurückgenommen, bleiben aber im Kontext. So lässt sich direkt ablesen, wo ein Open-Weight-Desktopmodell gegenüber einem proprietären Frontier-Modell steht, und ob der Leistungsunterschied den Preisunterschied rechtfertigt.
Im Scoreboard treten drei Lizenzmodelle gegeneinander an. Proprietäre Modelle geben weder Trainingsdaten heraus noch werden Gewichte veröffentlicht, der Zugang läuft ausschließlich über die API des Herstellers. Restricted-Weight-Modelle veröffentlichen ihre Gewichte und lassen sich lokal oder in der Cloud betreiben, allerdings schränkt die Lizenz die Nutzung meist auf akademische oder nicht-kommerzielle Zwecke ein. Open-Weight-Modelle stehen ohne Auflagen bereit: herunterladen, selbst betreiben, kommerziell einsetzen.
| # | Report | ||||||
|---|---|---|---|---|---|---|---|
|
1
|
GLM-5.3-Flash
Open Weight
Frontier
Cloud
|
80.6
|
Thinking
|
78Tool-Use unterstützt, Score: 78
|
52.7 t/s
|
$0.11
|
Report |
|
2
|
MiniMax M3
Open Weight
Frontier
Cloud
|
80.2
|
Adaptive
|
81Tool-Use unterstützt, Score: 81
|
108 t/s
|
$0.16
|
Report |
|
3
|
Qwen 3.8 27B (Thinking)
Open Weight
Workstation
vLLM
|
80.0
|
Thinking
|
87Tool-Use unterstützt, Score: 87
|
21.6 t/s
|
–
|
Report |
|
4
|
Claude Sonnet 4.6
Commercial
Frontier
API
|
79.8
|
Adaptive
|
75Tool-Use unterstützt, Score: 75
|
51.1 t/s
|
$1.77
|
Report |
|
5
|
Qwen3.8-Flash
Open Weight
Frontier
Cloud
|
79.5
|
Thinking
|
76Tool-Use unterstützt, Score: 76
|
50.9 t/s
|
$0.05
|
Report |
|
6
|
Claude Sonnet 5
Commercial
Frontier
API
|
79.0
|
Thinking
|
80Tool-Use unterstützt, Score: 80
|
82.5 t/s
|
$1.43
|
Report |
|
7
|
Xiaomi MiMo V2.5 Pro
Open Weight
Frontier
Cloud
|
79.0
|
Adaptive
|
74Tool-Use unterstützt, Score: 74
|
48.8 t/s
|
$0.12
|
Report |
|
8
|
Kimi K3
Open Weight
Frontier
Cloud
|
78.8
|
Thinking
|
81Tool-Use unterstützt, Score: 81
|
48.3 t/s
|
$3.30
|
Report |
|
9
|
Claude Opus 5
Commercial
Frontier
API
|
78.8
|
Thinking
|
80Tool-Use unterstützt, Score: 80
|
63.8 t/s
|
$4.62
|
Report |
|
10
|
Claude Opus 4.8
Commercial
Frontier
API
|
77.9
|
Thinking
|
72Tool-Use unterstützt, Score: 72
|
63.3 t/s
|
$2.71
|
Report |
|
11
|
Gemini 3.5 Flash
Commercial
Frontier
API
|
77.2
|
Adaptive
|
76Tool-Use unterstützt, Score: 76
|
58.7 t/s
|
$0.49
|
Report |
|
12
|
Meta Muse Spark 1.2
Commercial
Frontier
Cloud
|
76.7
|
Adaptive
|
70Tool-Use unterstützt, Score: 70
|
115 t/s
|
$0.54
|
Report |
|
13
|
Kimi K2.5
Open Weight
Frontier
Cloud
|
76.6
|
Thinking
|
77Tool-Use unterstützt, Score: 77
|
37.7 t/s
|
$0.39
|
Report |
|
14
|
GPT-5.5
Commercial
Frontier
API
|
76.5
|
Thinking
|
79Tool-Use unterstützt, Score: 79
|
59.3 t/s
|
$2.97
|
Report |
|
15
|
GLM 4.6
Restricted
Frontier
Cloud
|
76.2
|
Standard
|
53Tool-Use unterstützt, Score: 53
|
46.1 t/s
|
$0.25
|
Report |
|
16
|
Qwen 3.8 27B Uncensored (Thinking)
Open Weight
Workstation
vLLM
|
76.1
|
Thinking
|
74Tool-Use unterstützt, Score: 74
|
16.1 t/s
|
–
|
Report |
|
17
|
Ornith 1.5 35B-A3B
Open Weight
Workstation
vLLM
|
76.0
|
Thinking
|
76Tool-Use unterstützt, Score: 76
|
61.0 t/s
|
–
|
Report |
|
18
|
Qwen 3.7 Max
Commercial
Frontier
Cloud
|
75.8
|
Adaptive
|
60Tool-Use unterstützt, Score: 60
|
62.6 t/s
|
$0.75
|
Report |
|
19
|
Ornith 1.0 35B
Open Weight
Workstation
vLLM
|
75.8
|
Thinking
|
63Tool-Use unterstützt, Score: 63
|
35.5 t/s
|
–
|
Report |
|
20
|
Qwen 3.6 27B (Thinking)
Open Weight
Workstation
vLLM
|
75.6
|
Thinking
|
73Tool-Use unterstützt, Score: 73
|
25.2 t/s
|
–
|
Report |
|
21
|
Kimi K2.6
Open Weight
Frontier
Cloud
|
75.6
|
Adaptive
|
74Tool-Use unterstützt, Score: 74
|
49.4 t/s
|
$0.96
|
Report |
|
22
|
GLM-5.2
Open Weight
Frontier
Cloud
|
75.5
|
Adaptive
|
51Tool-Use unterstützt, Score: 51
|
65.7 t/s
|
$0.40
|
Report |
|
23
|
Qwen 3.6 Plus
Commercial
Frontier
Cloud
|
75.5
|
Adaptive
|
68Tool-Use unterstützt, Score: 68
|
56.3 t/s
|
$0.41
|
Report |
|
24
|
Upstage Solar Pro4
Commercial
Frontier
Cloud
|
75.4
|
Standard
|
71Tool-Use unterstützt, Score: 71
|
35.9 t/s
|
$0.01
|
Report |
|
25
|
Qwen 3.8 27B
Open Weight
Workstation
vLLM
|
75.2
|
Standard
|
70Tool-Use unterstützt, Score: 70
|
17.4 t/s
|
–
|
Report |
|
26
|
Gemma 4 26B-A4B Instruct (Thinking)
Open Weight
Workstation
vLLM
|
74.9
|
Thinking
|
74Tool-Use unterstützt, Score: 74
|
27.3 t/s
|
–
|
Report |
|
27
|
Claude Haiku 4.5
Commercial
Frontier
API
|
74.8
|
Standard
|
67Tool-Use unterstützt, Score: 67
|
98.4 t/s
|
$0.43
|
Report |
|
28
|
Mistral 3 Large
Open Weight
Frontier
API
|
74.8
|
Standard
|
69Tool-Use unterstützt, Score: 69
|
50.7 t/s
|
$0.10
|
Report |
|
29
|
Grok 4.6
Commercial
Frontier
API
|
74.7
|
Thinking
|
73Tool-Use unterstützt, Score: 73
|
13.6 t/s
|
$0.44
|
Report |
|
30
|
Xiaomi MiMo V2.5
Open Weight
Frontier
Cloud
|
74.5
|
Adaptive
|
76Tool-Use unterstützt, Score: 76
|
64.5 t/s
|
$0.04
|
Report |
|
31
|
GPT-OSS 120B (Thinking)
Open Weight
Frontier
vLLM
|
74.4
|
Thinking
|
76Tool-Use unterstützt, Score: 76
|
31.4 t/s
|
–
|
Report |
|
32
|
Qwen 3.5 35B-A3B (Unsloth)
Open Weight
Workstation
llama.cpp
|
74.3
|
Adaptive
|
72Tool-Use unterstützt, Score: 72
|
69.3 t/s
|
–
|
Report |
|
33
|
Mistral Medium 3.5
Open Weight
Frontier
API
|
74.2
|
Standard
|
71Tool-Use unterstützt, Score: 71
|
106 t/s
|
$0.46
|
Report |
|
34
|
Qwen 3 Coder Next
Open Weight
Frontier
llama.cpp
|
74.1
|
Standard
|
71Tool-Use unterstützt, Score: 71
|
48.5 t/s
|
–
|
Report |
|
35
|
Qwen 3.6 35B-A3B (Unsloth) (Thinking)
Open Weight
Workstation
vLLM
|
74.1
|
Thinking
|
77Tool-Use unterstützt, Score: 77
|
84.6 t/s
|
–
|
Report |
|
36
|
Grok 4.5
Commercial
Frontier
API
|
73.8
|
Standard
|
77Tool-Use unterstützt, Score: 77
|
25.6 t/s
|
$0.52
|
Report |
|
37
|
Qwen 3.6 27B
Open Weight
Workstation
vLLM
|
73.5
|
Standard
|
76Tool-Use unterstützt, Score: 76
|
19.0 t/s
|
–
|
Report |
|
38
|
Qwen 3.8 27B Uncensored
Open Weight
Workstation
vLLM
|
73.4
|
Standard
|
77Tool-Use unterstützt, Score: 77
|
13.4 t/s
|
–
|
Report |
|
39
|
Gemma 4 26B-A4B Q5_K_M (ARA-Abliterated)
Open Weight
Workstation
llama.cpp
|
73.3
|
Standard
|
76Tool-Use unterstützt, Score: 76
|
53.5 t/s
|
–
|
Report |
|
40
|
GLM-5.1
Open Weight
Frontier
Cloud
|
73.3
|
Adaptive
|
68Tool-Use unterstützt, Score: 68
|
63.5 t/s
|
$0.52
|
Report |
|
41
|
Mistral Small 4
Open Weight
Frontier
API
|
73.3
|
Standard
|
20Tool-Use unterstützt, Score: 20
|
117 t/s
|
$0.03
|
Report |
|
42
|
Gemma 4 31B Instruct
Open Weight
Workstation
Cloud
|
73.2
|
Standard
|
74Tool-Use unterstützt, Score: 74
|
30.6 t/s
|
$0.02
|
Report |
|
43
|
Qwen 3.5 397B A17B
Open Weight
Frontier
Cloud
|
73.2
|
Adaptive
|
74Tool-Use unterstützt, Score: 74
|
20.4 t/s
|
$0.46
|
Report |
|
44
|
Gemini 2.5 Pro
Commercial
Frontier
API
|
73.2
|
Adaptive
|
71Tool-Use unterstützt, Score: 71
|
29.7 t/s
|
$0.68
|
Report |
|
45
|
MiniMax M2.7
Restricted
Frontier
Cloud
|
73.0
|
Adaptive
|
65Tool-Use unterstützt, Score: 65
|
48.0 t/s
|
$0.13
|
Report |
|
46
|
Kimi K2.7 Code
Open Weight
Frontier
Cloud
|
73.0
|
Thinking
|
66Tool-Use unterstützt, Score: 66
|
100 t/s
|
$0.58
|
Report |
|
47
|
GPT-5.4
Commercial
Frontier
API
|
72.9
|
Standard
|
58Tool-Use unterstützt, Score: 58
|
63.6 t/s
|
$0.94
|
Report |
|
48
|
GLM-4.7
Restricted
Frontier
Cloud
|
72.9
|
Adaptive
|
71Tool-Use unterstützt, Score: 71
|
58.8 t/s
|
$0.29
|
Report |
|
49
|
Gemma 4 12B Instruct (Unsloth, Q6_K_XL) — Instruct-Profil
Open Weight
Desktop
llama.cpp
|
72.9
|
Standard
|
75Tool-Use unterstützt, Score: 75
|
13.3 t/s
|
–
|
Report |
|
50
|
DeepSeek V4 Flash
Open Weight
Frontier
Cloud
|
72.9
|
Adaptive
|
78Tool-Use unterstützt, Score: 78
|
29.9 t/s
|
$0.01
|
Report |
|
51
|
GPT-OSS 120B
Open Weight
Frontier
vLLM
|
72.8
|
Standard
|
67Tool-Use unterstützt, Score: 67
|
30.1 t/s
|
–
|
Report |
|
52
|
Gemma 4 12B Instruct (Unsloth, Q6_K_XL)
Open Weight
Desktop
llama.cpp
|
72.7
|
Thinking
|
71Tool-Use unterstützt, Score: 71
|
16.6 t/s
|
–
|
Report |
|
53
|
DeepSeek V4 Pro
Open Weight
Frontier
Cloud
|
72.6
|
Adaptive
|
77Tool-Use unterstützt, Score: 77
|
53.3 t/s
|
$0.30
|
Report |
|
54
|
NVIDIA Nemotron 3 Ultra 550B A55B
Open Weight
Frontier
Cloud
|
72.5
|
Adaptive
|
55Tool-Use unterstützt, Score: 55
|
97.0 t/s
|
$0.37
|
Report |
|
55
|
Gemma 4 26B-A4B Instruct
Open Weight
Workstation
vLLM
|
72.0
|
Standard
|
65Tool-Use unterstützt, Score: 65
|
35.2 t/s
|
–
|
Report |
|
56
|
Gemma 4 E4B
Open Weight
Edge
llama.cpp
|
71.9
|
Adaptive
|
73Tool-Use unterstützt, Score: 73
|
48.2 t/s
|
–
|
Report |
|
57
|
Qwen 3.6 35B-A3B (Uncensored)
Open Weight
Workstation
llama.cpp
|
71.7
|
Adaptive
|
60Tool-Use unterstützt, Score: 60
|
53.9 t/s
|
–
|
Report |
|
58
|
Gemini 3.7 Flash
Commercial
Frontier
Cloud
|
71.6
|
Adaptive
|
80Tool-Use unterstützt, Score: 80
|
94.1 t/s
|
–
|
Report |
|
59
|
Muse Glimmer 30B
Open Weight
Workstation
vLLM
|
71.6
|
Thinking
|
73Tool-Use unterstützt, Score: 73
|
10.6 t/s
|
–
|
Report |
|
60
|
Qwen 3.5 27B
Open Weight
Workstation
vLLM
|
71.0
|
Thinking
|
66Tool-Use unterstützt, Score: 66
|
17.8 t/s
|
–
|
Report |
|
61
|
Gemma 4 12B Instruct (Unsloth)
Open Weight
Desktop
llama.cpp
|
70.7
|
Adaptive
|
62Tool-Use unterstützt, Score: 62
|
13.1 t/s
|
–
|
Report |
|
62
|
Qwen 3.6 35B-A3B (Unsloth)
Open Weight
Workstation
vLLM
|
70.5
|
Standard
|
56Tool-Use unterstützt, Score: 56
|
67.5 t/s
|
–
|
Report |
|
63
|
Devstral 2
Open Weight
Frontier
API
|
70.2
|
Standard
|
55Tool-Use unterstützt, Score: 55
|
56.2 t/s
|
$0.13
|
Report |
|
64
|
GPT-5.4 Mini
Commercial
Frontier
API
|
70.2
|
Standard
|
68Tool-Use unterstützt, Score: 68
|
120 t/s
|
$0.25
|
Report |
|
65
|
Qwen 3.5 4B (Unsloth)
Open Weight
Nano
llama.cpp
|
69.9
|
Adaptive
|
78Tool-Use unterstützt, Score: 78
|
51.3 t/s
|
–
|
Report |
|
66
|
NVIDIA Nemotron 3.5 Lightning 30B (Thinking)
Open Weight
Workstation
vLLM
|
69.6
|
Thinking
|
73Tool-Use unterstützt, Score: 73
|
93.5 t/s
|
–
|
Report |
|
67
|
o4-mini
Commercial
Frontier
API
|
69.5
|
Thinking
|
63Tool-Use unterstützt, Score: 63
|
69.7 t/s
|
$0.40
|
Report |
|
68
|
Laguna S 2.1
Open Weight
Frontier
vLLM
|
69.1
|
Thinking
|
63Tool-Use unterstützt, Score: 63
|
17.2 t/s
|
$0.02
|
Report |
|
69
|
Hermes 4 70B
Open Weight
Server
Cloud
|
69.0
|
Adaptive
|
76Tool-Use unterstützt, Score: 76
|
76.8 t/s
|
$0.03
|
Report |
|
70
|
Gemma 3 12B IT
Restricted
Desktop
llama.cpp
|
68.9
|
Standard
|
64Tool-Use unterstützt, Score: 64
|
39.1 t/s
|
–
|
Report |
|
71
|
Qwen 3.5 9B (Unsloth)
Open Weight
Edge
llama.cpp
|
68.7
|
Adaptive
|
57Tool-Use unterstützt, Score: 57
|
34.3 t/s
|
–
|
Report |
|
72
|
Hermes 4 405B
Restricted
Frontier
Cloud
|
68.4
|
Adaptive
|
69Tool-Use unterstützt, Score: 69
|
31.1 t/s
|
$0.17
|
Report |
|
73
|
NVIDIA Nemotron 3 Nano 30B A3B
Open Weight
Workstation
Cloud
|
68.3
|
Adaptive
|
66Tool-Use unterstützt, Score: 66
|
35.4 t/s
|
$0.02
|
Report |
|
74
|
Hermes 4 14B (Abliterated)
Open Weight
Desktop
llama.cpp
|
68.2
|
Adaptive
|
70Tool-Use unterstützt, Score: 70
|
25.4 t/s
|
–
|
Report |
|
75
|
GPT-5.4 Nano
Commercial
Frontier
API
|
68.1
|
Standard
|
49Tool-Use unterstützt, Score: 49
|
125 t/s
|
$0.07
|
Report |
|
76
|
o3-mini
Commercial
Frontier
API
|
68.0
|
Thinking
|
73Tool-Use unterstützt, Score: 73
|
71.2 t/s
|
$0.37
|
Report |
|
77
|
Hermes 4.3 36B (Thinking)
Open Weight
Server
vLLM
|
68.0
|
Thinking
|
67Tool-Use unterstützt, Score: 67
|
13.3 t/s
|
–
|
Report |
|
78
|
GLM-5.3
Open Weight
Frontier
Cloud
|
67.7
|
Thinking
|
80Tool-Use unterstützt, Score: 80
|
53.7 t/s
|
$1.47
|
Report |
|
79
|
Qwen 3 14B
Open Weight
Desktop
llama.cpp
|
67.7
|
Adaptive
|
66Tool-Use unterstützt, Score: 66
|
23.9 t/s
|
–
|
Report |
|
80
|
Hermes 4 14B
Open Weight
Desktop
llama.cpp
|
67.3
|
Adaptive
|
67Tool-Use unterstützt, Score: 67
|
30.3 t/s
|
–
|
Report |
|
81
|
Qwen3.8-2.4T-A95B
Open Weight
Frontier
Cloud
|
67.2
|
Thinking
|
78Tool-Use unterstützt, Score: 78
|
44.2 t/s
|
$1.89
|
Report |
|
82
|
Hermes 4.3 36B
Open Weight
Server
vLLM
|
67.1
|
Standard
|
67Tool-Use unterstützt, Score: 67
|
12.7 t/s
|
–
|
Report |
|
83
|
Codestral 25.08
Restricted
Desktop
API
|
65.5
|
Standard
|
69Tool-Use unterstützt, Score: 69
|
192 t/s
|
$0.05
|
Report |
|
84
|
GPT-OSS 20B
Open Weight
Desktop
vLLM
|
64.3
|
Standard
|
28Tool-Use unterstützt, Score: 28
|
41.5 t/s
|
–
|
Report |
|
85
|
GPT-OSS 20B (Thinking)
Open Weight
Desktop
vLLM
|
61.6
|
Thinking
|
26Tool-Use unterstützt, Score: 26
|
41.2 t/s
|
–
|
Report |
|
86
|
Qwen 3 4B
Open Weight
Nano
llama.cpp
|
60.9
|
Adaptive
|
71Tool-Use unterstützt, Score: 71
|
74.0 t/s
|
–
|
Report |
|
87
|
Hermes 3 8B
Restricted
Edge
llama.cpp
|
58.8
|
Standard
|
56Tool-Use unterstützt, Score: 56
|
48.0 t/s
|
–
|
Report |
|
88
|
Qwen 2.5 Coder 7B
Open Weight
Edge
llama.cpp
|
55.3
|
Standard
|
58Tool-Use unterstützt, Score: 58
|
37.4 t/s
|
–
|
Report |