LLMs im direkten Vergleich 7 Einsatzszenarien · Gleiche Bedingungen · eine Ergebnisliste
Wie gut schlagen sich freie Open-Weight-Modelle gegen proprietäre API-Modelle und lizenzlich eingeschränkte Restricted-Weight-Modelle? Anhand alltäglicher Aufgaben und unter denselben Bedingungen landen alle Lizenzmodelle und Gewichtsklassen im selben Scoreboard.
Die Stärke des Scoreboards liegt in der Vergleichbarkeit. Wer nach Größenklasse oder Modelltyp filtert, blendet nichts aus. Nicht passende Einträge werden visuell zurückgenommen, bleiben aber im Kontext. So lässt sich direkt ablesen, wo ein Open-Weight-Desktopmodell gegenüber einem proprietären Frontier-Modell steht, und ob der Leistungsunterschied den Preisunterschied rechtfertigt.
Im Scoreboard treten drei Lizenzmodelle gegeneinander an. Proprietäre Modelle geben weder Trainingsdaten heraus noch werden Gewichte veröffentlicht, der Zugang läuft ausschließlich über die API des Herstellers. Restricted-Weight-Modelle veröffentlichen ihre Gewichte und lassen sich lokal oder in der Cloud betreiben, allerdings schränkt die Lizenz die Nutzung meist auf akademische oder nicht-kommerzielle Zwecke ein. Open-Weight-Modelle stehen ohne Auflagen bereit: herunterladen, selbst betreiben, kommerziell einsetzen.
| # | Report | ||||||
|---|---|---|---|---|---|---|---|
|
1
|
Claude Opus 4.7
Commercial
Frontier
API
|
79.3
|
Adaptive
|
83Tool-Use unterstützt, Score: 83
|
53.1 t/s
|
$2.73
|
Report |
|
2
|
Qwen 3.7 Max
Commercial
Frontier
LCL
|
78.8
|
Adaptive
|
70Tool-Use unterstützt, Score: 70
|
19.1 t/s
|
$0.64
|
Report |
|
3
|
Kimi K3
Open Weight
Frontier
LCL
|
78.0
|
Thinking
|
79Tool-Use unterstützt, Score: 79
|
16.1 t/s
|
$3.15
|
Report |
|
4
|
Claude Sonnet 4.6
Commercial
Frontier
API
|
78.0
|
Adaptive
|
66Tool-Use unterstützt, Score: 66
|
43.4 t/s
|
$1.76
|
Report |
|
5
|
MiniMax M3
Open Weight
Frontier
LCL
|
77.8
|
Adaptive
|
76Tool-Use unterstützt, Score: 76
|
45.3 t/s
|
$0.15
|
Report |
|
6
|
Claude Opus 4.8
Commercial
Frontier
API
|
77.5
|
Thinking
|
71Tool-Use unterstützt, Score: 71
|
63.2 t/s
|
$2.66
|
Report |
|
7
|
Xiaomi MiMo V2.5 Pro
Open Weight
Frontier
LCL
|
77.0
|
Adaptive
|
68Tool-Use unterstützt, Score: 68
|
27.4 t/s
|
$0.12
|
Report |
|
8
|
Claude Sonnet 5
Commercial
Frontier
API
|
76.9
|
Thinking
|
74Tool-Use unterstützt, Score: 74
|
64.5 t/s
|
$1.41
|
Report |
|
9
|
Ornith 1.0 35B (FP8) (Thinking)
Open Weight
Workstation
LCL
|
76.9
|
Thinking
|
74Tool-Use unterstützt, Score: 74
|
21.6 t/s
|
–
|
Report |
|
10
|
Kimi K2.7 Code
Open Weight
Frontier
LCL
|
76.4
|
Thinking
|
71Tool-Use unterstützt, Score: 71
|
70.8 t/s
|
$0.61
|
Report |
|
11
|
GLM-5.1
Open Weight
Frontier
LCL
|
76.2
|
Adaptive
|
74Tool-Use unterstützt, Score: 74
|
27.0 t/s
|
$0.43
|
Report |
|
12
|
Claude Opus 4.6
Commercial
Frontier
API
|
76.1
|
Adaptive
|
66Tool-Use unterstützt, Score: 66
|
44.0 t/s
|
$2.78
|
Report |
|
13
|
Qwen 3.6 Plus
Commercial
Frontier
LCL
|
76.0
|
Adaptive
|
59Tool-Use unterstützt, Score: 59
|
18.3 t/s
|
$0.35
|
Report |
|
14
|
Kimi K2.6
Open Weight
Frontier
LCL
|
75.9
|
Adaptive
|
75Tool-Use unterstützt, Score: 75
|
25.1 t/s
|
$0.75
|
Report |
|
15
|
Gemini 2.5 Pro
Commercial
Frontier
API
|
75.3
|
Adaptive
|
74Tool-Use unterstützt, Score: 74
|
30.9 t/s
|
$0.69
|
Report |
|
16
|
MiniMax M2.7
Restricted
Frontier
LCL
|
75.3
|
Adaptive
|
51Tool-Use unterstützt, Score: 51
|
69.0 t/s
|
$0.11
|
Report |
|
17
|
Gemma 4 31B Instruct
Open Weight
Workstation
LCL
|
75.2
|
Standard
|
69Tool-Use unterstützt, Score: 69
|
16.9 t/s
|
–
|
Report |
|
18
|
Claude Haiku 4.5
Commercial
Frontier
API
|
75.1
|
Standard
|
66Tool-Use unterstützt, Score: 66
|
100 t/s
|
$0.45
|
Report |
|
19
|
DeepSeek V4 Pro
Open Weight
Frontier
LCL
|
75.1
|
Adaptive
|
76Tool-Use unterstützt, Score: 76
|
36.3 t/s
|
$0.09
|
Report |
|
20
|
Kimi K2 Thinking
Open Weight
Frontier
LCL
|
75.0
|
Thinking
|
76Tool-Use unterstützt, Score: 76
|
34.9 t/s
|
$0.37
|
Report |
|
21
|
Claude Sonnet 4.5
Commercial
Frontier
API
|
74.9
|
Adaptive
|
72Tool-Use unterstützt, Score: 72
|
50.0 t/s
|
$1.24
|
Report |
|
22
|
GLM-5
Open Weight
Frontier
LCL
|
74.8
|
Adaptive
|
74Tool-Use unterstützt, Score: 74
|
39.8 t/s
|
$0.25
|
Report |
|
23
|
DeepSeek V4 Flash
Open Weight
Frontier
LCL
|
74.6
|
Adaptive
|
78Tool-Use unterstützt, Score: 78
|
44.0 t/s
|
$0.02
|
Report |
|
24
|
GPT-5
Commercial
Frontier
API
|
74.5
|
Adaptive
|
62Tool-Use unterstützt, Score: 62
|
34.9 t/s
|
$1.75
|
Report |
|
25
|
GLM-4.7
Restricted
Frontier
LCL
|
74.5
|
Adaptive
|
63Tool-Use unterstützt, Score: 63
|
26.1 t/s
|
$0.25
|
Report |
|
26
|
GLM-5.2
Open Weight
Frontier
LCL
|
74.3
|
Adaptive
|
76Tool-Use unterstützt, Score: 76
|
37.6 t/s
|
$0.48
|
Report |
|
27
|
GPT-5.5
Commercial
Frontier
API
|
74.3
|
Thinking
|
71Tool-Use unterstützt, Score: 71
|
38.6 t/s
|
$2.99
|
Report |
|
28
|
Qwen 3.5 35B-A3B (Unsloth)
Open Weight
Workstation
LCL
|
74.3
|
Adaptive
|
72Tool-Use unterstützt, Score: 72
|
69.3 t/s
|
–
|
Report |
|
29
|
Qwen 3 Coder Next
Open Weight
Workstation
LCL
|
74.1
|
Standard
|
71Tool-Use unterstützt, Score: 71
|
48.5 t/s
|
–
|
Report |
|
30
|
GPT-OSS 120B
Open Weight
Workstation
LCL
|
74.1
|
Adaptive
|
69Tool-Use unterstützt, Score: 69
|
369 t/s
|
$0.05
|
Report |
|
31
|
Gemma 4 26B-A4B Instruct
Open Weight
Workstation
LCL
|
74.0
|
Standard
|
72Tool-Use unterstützt, Score: 72
|
47.8 t/s
|
–
|
Report |
|
32
|
Qwen3.6 27B Instruct (Thinking)
Open Weight
Workstation
LCL
|
73.9
|
Thinking
|
20Tool-Use unterstützt, Score: 20
|
7.2 t/s
|
–
|
Report |
|
33
|
Mistral Medium 3.5
Open Weight
Frontier
API
|
73.9
|
Standard
|
74Tool-Use unterstützt, Score: 74
|
135 t/s
|
$0.48
|
Report |
|
34
|
Gemma 4 Ortenzya Creative Wordsmith 31B
Open Weight
Workstation
LCL
|
73.8
|
Standard
|
75Tool-Use unterstützt, Score: 75
|
17.5 t/s
|
–
|
Report |
|
35
|
Gemini 3.5 Flash
Commercial
Frontier
API
|
73.7
|
Adaptive
|
72Tool-Use unterstützt, Score: 72
|
51.9 t/s
|
$0.47
|
Report |
|
36
|
Ornith 1.0 35B (FP8)
Open Weight
Workstation
LCL
|
73.4
|
Standard
|
71Tool-Use unterstützt, Score: 71
|
49.3 t/s
|
–
|
Report |
|
37
|
Qwen 3.6 27B NVFP4 (vLLM, Dense, MTP)
Open Weight
Workstation
LCL
|
73.4
|
Standard
|
65Tool-Use unterstützt, Score: 65
|
22.6 t/s
|
–
|
Report |
|
38
|
Gemma 4 ARA 26B-A4B (ARA-Abliterated)
Open Weight
Workstation
LCL
|
73.3
|
Standard
|
76Tool-Use unterstützt, Score: 76
|
53.5 t/s
|
–
|
Report |
|
39
|
Kimi K2.5
Open Weight
Frontier
LCL
|
73.3
|
Thinking
|
77Tool-Use unterstützt, Score: 77
|
21.1 t/s
|
$0.33
|
Report |
|
40
|
Gemma 4 31B Instruct
Open Weight
Workstation
LCL
|
73.2
|
Standard
|
74Tool-Use unterstützt, Score: 74
|
31.2 t/s
|
$0.02
|
Report |
|
41
|
Qwen 3.6 35B-A3B NVFP4 (vLLM, MoE, MTP) (Thinking)
Open Weight
Workstation
LCL
|
73.2
|
Thinking
|
68Tool-Use unterstützt, Score: 68
|
36.6 t/s
|
–
|
Report |
|
42
|
Qwen 3.5 397B A17B
Open Weight
Frontier
LCL
|
73.2
|
Adaptive
|
74Tool-Use unterstützt, Score: 74
|
20.4 t/s
|
$0.46
|
Report |
|
43
|
Mistral 3 Large
Open Weight
Frontier
API
|
73.2
|
Standard
|
60Tool-Use unterstützt, Score: 60
|
55.1 t/s
|
$0.43
|
Report |
|
44
|
GLM 4.6
Restricted
Frontier
LCL
|
73.1
|
Standard
|
76Tool-Use unterstützt, Score: 76
|
17.8 t/s
|
$0.27
|
Report |
|
45
|
GPT-5.4
Commercial
Frontier
API
|
72.8
|
Standard
|
57Tool-Use unterstützt, Score: 57
|
78.7 t/s
|
$0.98
|
Report |
|
46
|
GLM-5 Turbo
Commercial
Frontier
LCL
|
72.7
|
Adaptive
|
79Tool-Use unterstützt, Score: 79
|
20.9 t/s
|
$0.45
|
Report |
|
47
|
NVIDIA Nemotron 3 Ultra 550B A55B
Open Weight
Frontier
LCL
|
72.5
|
Adaptive
|
55Tool-Use unterstützt, Score: 55
|
98.9 t/s
|
$0.26
|
Report |
|
48
|
Qwen 3.6 35B-A3B (Unsloth)
Open Weight
Desktop
LCL
|
72.5
|
Adaptive
|
66Tool-Use unterstützt, Score: 66
|
66.4 t/s
|
–
|
Report |
|
49
|
GPT-5 Mini
Commercial
Nano
API
|
72.3
|
Standard
|
17Tool-Use unterstützt, Score: 17
|
34.6 t/s
|
$0.23
|
Report |
|
50
|
Grok 4.5
Commercial
Frontier
API
|
72.1
|
Standard
|
74Tool-Use unterstützt, Score: 74
|
46.7 t/s
|
$0.42
|
Report |
|
51
|
DeepSeek V3.2
Open Weight
Frontier
LCL
|
71.9
|
Standard
|
67Tool-Use unterstützt, Score: 67
|
46.3 t/s
|
$0.02
|
Report |
|
52
|
Gemma 4 Ortenzya Creative Wordsmith 31B (Thinking)
Open Weight
Workstation
LCL
|
71.9
|
Thinking
|
66Tool-Use unterstützt, Score: 66
|
7.1 t/s
|
–
|
Report |
|
53
|
Gemma 4 E4B
Open Weight
Edge
LCL
|
71.9
|
Adaptive
|
73Tool-Use unterstützt, Score: 73
|
48.2 t/s
|
–
|
Report |
|
54
|
Qwen 3.6 27B NVFP4 (vLLM, Dense, MTP) (Thinking)
Open Weight
Workstation
LCL
|
71.7
|
Thinking
|
60Tool-Use unterstützt, Score: 60
|
8.5 t/s
|
–
|
Report |
|
55
|
Qwen 3.6 35B-A3B (Uncensored)
Open Weight
Desktop
LCL
|
71.7
|
Adaptive
|
60Tool-Use unterstützt, Score: 60
|
53.9 t/s
|
–
|
Report |
|
56
|
Gemma 4 31B Instruct (Thinking)
Open Weight
Workstation
LCL
|
71.6
|
Thinking
|
69Tool-Use unterstützt, Score: 69
|
7.5 t/s
|
–
|
Report |
|
57
|
Kimi K2
Open Weight
Frontier
LCL
|
71.3
|
Standard
|
71Tool-Use unterstützt, Score: 71
|
21.3 t/s
|
$0.12
|
Report |
|
58
|
Qwen3.6 27B Instruct
Open Weight
Workstation
LCL
|
71.1
|
Standard
|
64Tool-Use unterstützt, Score: 64
|
22.1 t/s
|
–
|
Report |
|
59
|
Gemma 4 12B Instruct (Unsloth)
Open Weight
Desktop
LCL
|
70.7
|
Adaptive
|
62Tool-Use unterstützt, Score: 62
|
13.1 t/s
|
–
|
Report |
|
60
|
Mistral Small 4
Open Weight
Workstation
API
|
70.7
|
Standard
|
53Tool-Use unterstützt, Score: 53
|
177 t/s
|
$0.02
|
Report |
|
61
|
Xiaomi MiMo V2.5
Open Weight
Frontier
LCL
|
70.6
|
Adaptive
|
79Tool-Use unterstützt, Score: 79
|
37.5 t/s
|
$0.23
|
Report |
|
62
|
Hermes 4 70B
Open Weight
Server
LCL
|
70.4
|
Adaptive
|
66Tool-Use unterstützt, Score: 66
|
75.8 t/s
|
$0.02
|
Report |
|
63
|
Qwen 3.6 35B-A3B NVFP4 (vLLM, MoE, MTP)
Open Weight
Workstation
LCL
|
70.3
|
Standard
|
63Tool-Use unterstützt, Score: 63
|
79.1 t/s
|
–
|
Report |
|
64
|
Gemma 4 26B-A4B Instruct (Thinking)
Open Weight
Workstation
LCL
|
70.3
|
Thinking
|
69Tool-Use unterstützt, Score: 69
|
18.6 t/s
|
–
|
Report |
|
65
|
Llama 3.3 Nemotron Super 49B v1.5
Open Weight
Server
LCL
|
70.2
|
Adaptive
|
73Tool-Use unterstützt, Score: 73
|
20.7 t/s
|
$0.04
|
Report |
|
66
|
Gemma 4 31B Creative Wordsmith (Uncensored)
Open Weight
Workstation
LCL
|
70.2
|
Adaptive
|
66Tool-Use unterstützt, Score: 66
|
9.4 t/s
|
–
|
Report |
|
67
|
Devstral 2
Open Weight
Frontier
API
|
70.2
|
Standard
|
55Tool-Use unterstützt, Score: 55
|
56.2 t/s
|
$0.12
|
Report |
|
68
|
GPT-5.4 Mini
Commercial
Frontier
API
|
70.2
|
Standard
|
68Tool-Use unterstützt, Score: 68
|
120 t/s
|
$0.25
|
Report |
|
69
|
Grok 4 (Non-Reasoning)
Commercial
Frontier
API
|
70.1
|
Adaptive
|
55Tool-Use unterstützt, Score: 55
|
182 t/s
|
$0.17
|
Report |
|
70
|
Grok 4.20 (Reasoning)
Commercial
Frontier
API
|
70.1
|
Thinking
|
75Tool-Use unterstützt, Score: 75
|
41.0 t/s
|
$0.15
|
Report |
|
71
|
Qwen 3.5 4B (Unsloth)
Open Weight
Nano
LCL
|
69.9
|
Adaptive
|
78Tool-Use unterstützt, Score: 78
|
51.3 t/s
|
–
|
Report |
|
72
|
o4-mini
Commercial
Frontier
API
|
69.5
|
Thinking
|
63Tool-Use unterstützt, Score: 63
|
69.7 t/s
|
$0.40
|
Report |
|
73
|
Qwen 3 32B
Open Weight
Workstation
LCL
|
69.3
|
Adaptive
|
65Tool-Use unterstützt, Score: 65
|
161 t/s
|
$0.06
|
Report |
|
74
|
Gemma 3 12B IT
Restricted
Desktop
LCL
|
68.9
|
Standard
|
64Tool-Use unterstützt, Score: 64
|
39.1 t/s
|
–
|
Report |
|
75
|
GPT-5.4 Nano
Commercial
Frontier
API
|
68.9
|
Standard
|
58Tool-Use unterstützt, Score: 58
|
123 t/s
|
$0.07
|
Report |
|
76
|
Qwen 3.5 9B (Unsloth)
Open Weight
Edge
LCL
|
68.7
|
Adaptive
|
57Tool-Use unterstützt, Score: 57
|
34.3 t/s
|
–
|
Report |
|
77
|
NVIDIA Nemotron 3 Nano 30B A3B
Open Weight
Workstation
LCL
|
68.3
|
Adaptive
|
66Tool-Use unterstützt, Score: 66
|
36.1 t/s
|
$0.02
|
Report |
|
78
|
Hermes 4 14B (Abliterated)
Open Weight
Desktop
LCL
|
68.2
|
Adaptive
|
70Tool-Use unterstützt, Score: 70
|
25.4 t/s
|
–
|
Report |
|
79
|
o3-mini
Commercial
Frontier
API
|
68.0
|
Thinking
|
73Tool-Use unterstützt, Score: 73
|
71.2 t/s
|
$0.37
|
Report |
|
80
|
GPT-4o
Commercial
Frontier
API
|
68.0
|
Standard
|
67Tool-Use unterstützt, Score: 67
|
150 t/s
|
$0.48
|
Report |
|
81
|
Hermes 4 405B
Restricted
Frontier
LCL
|
67.8
|
Adaptive
|
74Tool-Use unterstützt, Score: 74
|
39.5 t/s
|
$0.15
|
Report |
|
82
|
Qwen 3 14B
Open Weight
Desktop
LCL
|
67.7
|
Adaptive
|
66Tool-Use unterstützt, Score: 66
|
23.9 t/s
|
–
|
Report |
|
83
|
OpenAI o1
Commercial
Frontier
API
|
67.5
|
Thinking
|
77Tool-Use unterstützt, Score: 77
|
52.5 t/s
|
$6.46
|
Report |
|
84
|
Hermes 4 14B
Open Weight
Desktop
LCL
|
67.3
|
Adaptive
|
67Tool-Use unterstützt, Score: 67
|
30.3 t/s
|
–
|
Report |
|
85
|
Grok 4.3
Commercial
Frontier
API
|
67.3
|
Thinking
|
62Tool-Use unterstützt, Score: 62
|
65.7 t/s
|
$0.12
|
Report |
|
86
|
Codestral 25.08
Restricted
Frontier
API
|
65.5
|
Standard
|
69Tool-Use unterstützt, Score: 69
|
192 t/s
|
$0.03
|
Report |
|
87
|
GPT-4o Mini
Commercial
Frontier
API
|
65.0
|
Standard
|
66Tool-Use unterstützt, Score: 66
|
80.8 t/s
|
$0.03
|
Report |
|
88
|
Llama 3.3 70B Versatile
Restricted
Server
LCL
|
64.3
|
Standard
|
42Tool-Use unterstützt, Score: 42
|
276 t/s
|
$0.03
|
Report |
|
89
|
Command A+
Open Weight
Frontier
LCL
|
61.9
|
Thinking
|
6Tool-Use unterstützt, Score: 6
|
76.5 t/s
|
–
|
Report |
|
90
|
Qwen 3 4B
Open Weight
Nano
LCL
|
60.9
|
Adaptive
|
71Tool-Use unterstützt, Score: 71
|
74.0 t/s
|
–
|
Report |
|
91
|
Hermes 3 8B
Restricted
Edge
LCL
|
58.8
|
Standard
|
56Tool-Use unterstützt, Score: 56
|
48.0 t/s
|
–
|
Report |
|
92
|
GPT-OSS 20B
Open Weight
Desktop
LCL
|
56.2
|
Adaptive
|
0Tool-Use unterstützt, Score: 0
|
406 t/s
|
$0.03
|
Report |
|
93
|
Qwen 2.5 Coder 7B
Open Weight
Nano
LCL
|
56.1
|
Standard
|
58Tool-Use unterstützt, Score: 58
|
50.1 t/s
|
–
|
Report |
|
94
|
Llama 4 Scout 17B
Restricted
Server
LCL
|
55.2
|
Standard
|
33Tool-Use unterstützt, Score: 33
|
355 t/s
|
$0.01
|
Report |