LLMs im direkten Vergleich 7 Einsatzszenarien · Gleiche Bedingungen · eine Ergebnisliste

Wie gut schlagen sich freie Open-Weight-Modelle gegen proprietäre API-Modelle und lizenzlich eingeschränkte Restricted-Weight-Modelle? Anhand alltäglicher Aufgaben und unter denselben Bedingungen landen alle Lizenzmodelle und Gewichtsklassen im selben Scoreboard.

Der Gesamtscore zeigt den besten Allrounder. Die Profile darunter beantworten die eigentliche Alltagsfrage: Reicht dieses Modell für deine Aufgabe – auch wenn es ein kleiner, lokaler Spezialist ist statt ein Cloud-Gigant?

Scoreboard

88 Modelle im Scoreboard

# Report
1
GLM-5.3-Flash
Open Weight Frontier Cloud
80.6
Thinking
78Tool-Use unterstützt, Score: 78
52.7 t/s
$0.11
Report
2
MiniMax M3
Open Weight Frontier Cloud
80.2
Adaptive
81Tool-Use unterstützt, Score: 81
108 t/s
$0.16
Report
3
Qwen 3.8 27B (Thinking)
Open Weight Workstation vLLM
80.0
Thinking
87Tool-Use unterstützt, Score: 87
21.6 t/s
Report
4
Claude Sonnet 4.6
Commercial Frontier API
79.8
Adaptive
75Tool-Use unterstützt, Score: 75
51.1 t/s
$1.77
Report
5
Qwen3.8-Flash
Open Weight Frontier Cloud
79.5
Thinking
76Tool-Use unterstützt, Score: 76
50.9 t/s
$0.05
Report
6
Claude Sonnet 5
Commercial Frontier API
79.0
Thinking
80Tool-Use unterstützt, Score: 80
82.5 t/s
$1.43
Report
7
Xiaomi MiMo V2.5 Pro
Open Weight Frontier Cloud
79.0
Adaptive
74Tool-Use unterstützt, Score: 74
48.8 t/s
$0.12
Report
8
Kimi K3
Open Weight Frontier Cloud
78.8
Thinking
81Tool-Use unterstützt, Score: 81
48.3 t/s
$3.30
Report
9
Claude Opus 5
Commercial Frontier API
78.8
Thinking
80Tool-Use unterstützt, Score: 80
63.8 t/s
$4.62
Report
10
Claude Opus 4.8
Commercial Frontier API
77.9
Thinking
72Tool-Use unterstützt, Score: 72
63.3 t/s
$2.71
Report
11
Gemini 3.5 Flash
Commercial Frontier API
77.2
Adaptive
76Tool-Use unterstützt, Score: 76
58.7 t/s
$0.49
Report
12
Meta Muse Spark 1.2
Commercial Frontier Cloud
76.7
Adaptive
70Tool-Use unterstützt, Score: 70
115 t/s
$0.54
Report
13
Kimi K2.5
Open Weight Frontier Cloud
76.6
Thinking
77Tool-Use unterstützt, Score: 77
37.7 t/s
$0.39
Report
14
GPT-5.5
Commercial Frontier API
76.5
Thinking
79Tool-Use unterstützt, Score: 79
59.3 t/s
$2.97
Report
15
GLM 4.6
Restricted Frontier Cloud
76.2
Standard
53Tool-Use unterstützt, Score: 53
46.1 t/s
$0.25
Report
16
Qwen 3.8 27B Uncensored (Thinking)
Open Weight Workstation vLLM
76.1
Thinking
74Tool-Use unterstützt, Score: 74
16.1 t/s
Report
17
Ornith 1.5 35B-A3B
Open Weight Workstation vLLM
76.0
Thinking
76Tool-Use unterstützt, Score: 76
61.0 t/s
Report
18
Qwen 3.7 Max
Commercial Frontier Cloud
75.8
Adaptive
60Tool-Use unterstützt, Score: 60
62.6 t/s
$0.75
Report
19
Ornith 1.0 35B
Open Weight Workstation vLLM
75.8
Thinking
63Tool-Use unterstützt, Score: 63
35.5 t/s
Report
20
Qwen 3.6 27B (Thinking)
Open Weight Workstation vLLM
75.6
Thinking
73Tool-Use unterstützt, Score: 73
25.2 t/s
Report
21
Kimi K2.6
Open Weight Frontier Cloud
75.6
Adaptive
74Tool-Use unterstützt, Score: 74
49.4 t/s
$0.96
Report
22
GLM-5.2
Open Weight Frontier Cloud
75.5
Adaptive
51Tool-Use unterstützt, Score: 51
65.7 t/s
$0.40
Report
23
Qwen 3.6 Plus
Commercial Frontier Cloud
75.5
Adaptive
68Tool-Use unterstützt, Score: 68
56.3 t/s
$0.41
Report
24
Upstage Solar Pro4
Commercial Frontier Cloud
75.4
Standard
71Tool-Use unterstützt, Score: 71
35.9 t/s
$0.01
Report
25
Qwen 3.8 27B
Open Weight Workstation vLLM
75.2
Standard
70Tool-Use unterstützt, Score: 70
17.4 t/s
Report
26
Gemma 4 26B-A4B Instruct (Thinking)
Open Weight Workstation vLLM
74.9
Thinking
74Tool-Use unterstützt, Score: 74
27.3 t/s
Report
27
Claude Haiku 4.5
Commercial Frontier API
74.8
Standard
67Tool-Use unterstützt, Score: 67
98.4 t/s
$0.43
Report
28
Mistral 3 Large
Open Weight Frontier API
74.8
Standard
69Tool-Use unterstützt, Score: 69
50.7 t/s
$0.10
Report
29
Grok 4.6
Commercial Frontier API
74.7
Thinking
73Tool-Use unterstützt, Score: 73
13.6 t/s
$0.44
Report
30
Xiaomi MiMo V2.5
Open Weight Frontier Cloud
74.5
Adaptive
76Tool-Use unterstützt, Score: 76
64.5 t/s
$0.04
Report
31
GPT-OSS 120B (Thinking)
Open Weight Frontier vLLM
74.4
Thinking
76Tool-Use unterstützt, Score: 76
31.4 t/s
Report
32
Qwen 3.5 35B-A3B (Unsloth)
Open Weight Workstation llama.cpp
74.3
Adaptive
72Tool-Use unterstützt, Score: 72
69.3 t/s
Report
33
Mistral Medium 3.5
Open Weight Frontier API
74.2
Standard
71Tool-Use unterstützt, Score: 71
106 t/s
$0.46
Report
34
Qwen 3 Coder Next
Open Weight Frontier llama.cpp
74.1
Standard
71Tool-Use unterstützt, Score: 71
48.5 t/s
Report
35
Qwen 3.6 35B-A3B (Unsloth) (Thinking)
Open Weight Workstation vLLM
74.1
Thinking
77Tool-Use unterstützt, Score: 77
84.6 t/s
Report
36
Grok 4.5
Commercial Frontier API
73.8
Standard
77Tool-Use unterstützt, Score: 77
25.6 t/s
$0.52
Report
37
Qwen 3.6 27B
Open Weight Workstation vLLM
73.5
Standard
76Tool-Use unterstützt, Score: 76
19.0 t/s
Report
38
Qwen 3.8 27B Uncensored
Open Weight Workstation vLLM
73.4
Standard
77Tool-Use unterstützt, Score: 77
13.4 t/s
Report
39
Gemma 4 26B-A4B Q5_K_M (ARA-Abliterated)
Open Weight Workstation llama.cpp
73.3
Standard
76Tool-Use unterstützt, Score: 76
53.5 t/s
Report
40
GLM-5.1
Open Weight Frontier Cloud
73.3
Adaptive
68Tool-Use unterstützt, Score: 68
63.5 t/s
$0.52
Report
41
Mistral Small 4
Open Weight Frontier API
73.3
Standard
20Tool-Use unterstützt, Score: 20
117 t/s
$0.03
Report
42
Gemma 4 31B Instruct
Open Weight Workstation Cloud
73.2
Standard
74Tool-Use unterstützt, Score: 74
30.6 t/s
$0.02
Report
43
Qwen 3.5 397B A17B
Open Weight Frontier Cloud
73.2
Adaptive
74Tool-Use unterstützt, Score: 74
20.4 t/s
$0.46
Report
44
Gemini 2.5 Pro
Commercial Frontier API
73.2
Adaptive
71Tool-Use unterstützt, Score: 71
29.7 t/s
$0.68
Report
45
MiniMax M2.7
Restricted Frontier Cloud
73.0
Adaptive
65Tool-Use unterstützt, Score: 65
48.0 t/s
$0.13
Report
46
Kimi K2.7 Code
Open Weight Frontier Cloud
73.0
Thinking
66Tool-Use unterstützt, Score: 66
100 t/s
$0.58
Report
47
GPT-5.4
Commercial Frontier API
72.9
Standard
58Tool-Use unterstützt, Score: 58
63.6 t/s
$0.94
Report
48
GLM-4.7
Restricted Frontier Cloud
72.9
Adaptive
71Tool-Use unterstützt, Score: 71
58.8 t/s
$0.29
Report
49
Gemma 4 12B Instruct (Unsloth, Q6_K_XL) — Instruct-Profil
Open Weight Desktop llama.cpp
72.9
Standard
75Tool-Use unterstützt, Score: 75
13.3 t/s
Report
50
DeepSeek V4 Flash
Open Weight Frontier Cloud
72.9
Adaptive
78Tool-Use unterstützt, Score: 78
29.9 t/s
$0.01
Report
51
GPT-OSS 120B
Open Weight Frontier vLLM
72.8
Standard
67Tool-Use unterstützt, Score: 67
30.1 t/s
Report
52
Gemma 4 12B Instruct (Unsloth, Q6_K_XL)
Open Weight Desktop llama.cpp
72.7
Thinking
71Tool-Use unterstützt, Score: 71
16.6 t/s
Report
53
DeepSeek V4 Pro
Open Weight Frontier Cloud
72.6
Adaptive
77Tool-Use unterstützt, Score: 77
53.3 t/s
$0.30
Report
54
NVIDIA Nemotron 3 Ultra 550B A55B
Open Weight Frontier Cloud
72.5
Adaptive
55Tool-Use unterstützt, Score: 55
97.0 t/s
$0.37
Report
55
Gemma 4 26B-A4B Instruct
Open Weight Workstation vLLM
72.0
Standard
65Tool-Use unterstützt, Score: 65
35.2 t/s
Report
56
Gemma 4 E4B
Open Weight Edge llama.cpp
71.9
Adaptive
73Tool-Use unterstützt, Score: 73
48.2 t/s
Report
57
Qwen 3.6 35B-A3B (Uncensored)
Open Weight Workstation llama.cpp
71.7
Adaptive
60Tool-Use unterstützt, Score: 60
53.9 t/s
Report
58
Gemini 3.7 Flash
Commercial Frontier Cloud
71.6
Adaptive
80Tool-Use unterstützt, Score: 80
94.1 t/s
Report
59
Muse Glimmer 30B
Open Weight Workstation vLLM
71.6
Thinking
73Tool-Use unterstützt, Score: 73
10.6 t/s
Report
60
Qwen 3.5 27B
Open Weight Workstation vLLM
71.0
Thinking
66Tool-Use unterstützt, Score: 66
17.8 t/s
Report
61
Gemma 4 12B Instruct (Unsloth)
Open Weight Desktop llama.cpp
70.7
Adaptive
62Tool-Use unterstützt, Score: 62
13.1 t/s
Report
62
Qwen 3.6 35B-A3B (Unsloth)
Open Weight Workstation vLLM
70.5
Standard
56Tool-Use unterstützt, Score: 56
67.5 t/s
Report
63
Devstral 2
Open Weight Frontier API
70.2
Standard
55Tool-Use unterstützt, Score: 55
56.2 t/s
$0.13
Report
64
GPT-5.4 Mini
Commercial Frontier API
70.2
Standard
68Tool-Use unterstützt, Score: 68
120 t/s
$0.25
Report
65
Qwen 3.5 4B (Unsloth)
Open Weight Nano llama.cpp
69.9
Adaptive
78Tool-Use unterstützt, Score: 78
51.3 t/s
Report
66
NVIDIA Nemotron 3.5 Lightning 30B (Thinking)
Open Weight Workstation vLLM
69.6
Thinking
73Tool-Use unterstützt, Score: 73
93.5 t/s
Report
67
o4-mini
Commercial Frontier API
69.5
Thinking
63Tool-Use unterstützt, Score: 63
69.7 t/s
$0.40
Report
68
Laguna S 2.1
Open Weight Frontier vLLM
69.1
Thinking
63Tool-Use unterstützt, Score: 63
17.2 t/s
$0.02
Report
69
Hermes 4 70B
Open Weight Server Cloud
69.0
Adaptive
76Tool-Use unterstützt, Score: 76
76.8 t/s
$0.03
Report
70
Gemma 3 12B IT
Restricted Desktop llama.cpp
68.9
Standard
64Tool-Use unterstützt, Score: 64
39.1 t/s
Report
71
Qwen 3.5 9B (Unsloth)
Open Weight Edge llama.cpp
68.7
Adaptive
57Tool-Use unterstützt, Score: 57
34.3 t/s
Report
72
Hermes 4 405B
Restricted Frontier Cloud
68.4
Adaptive
69Tool-Use unterstützt, Score: 69
31.1 t/s
$0.17
Report
73
NVIDIA Nemotron 3 Nano 30B A3B
Open Weight Workstation Cloud
68.3
Adaptive
66Tool-Use unterstützt, Score: 66
35.4 t/s
$0.02
Report
74
Hermes 4 14B (Abliterated)
Open Weight Desktop llama.cpp
68.2
Adaptive
70Tool-Use unterstützt, Score: 70
25.4 t/s
Report
75
GPT-5.4 Nano
Commercial Frontier API
68.1
Standard
49Tool-Use unterstützt, Score: 49
125 t/s
$0.07
Report
76
o3-mini
Commercial Frontier API
68.0
Thinking
73Tool-Use unterstützt, Score: 73
71.2 t/s
$0.37
Report
77
Hermes 4.3 36B (Thinking)
Open Weight Server vLLM
68.0
Thinking
67Tool-Use unterstützt, Score: 67
13.3 t/s
Report
78
GLM-5.3
Open Weight Frontier Cloud
67.7
Thinking
80Tool-Use unterstützt, Score: 80
53.7 t/s
$1.47
Report
79
Qwen 3 14B
Open Weight Desktop llama.cpp
67.7
Adaptive
66Tool-Use unterstützt, Score: 66
23.9 t/s
Report
80
Hermes 4 14B
Open Weight Desktop llama.cpp
67.3
Adaptive
67Tool-Use unterstützt, Score: 67
30.3 t/s
Report
81
Qwen3.8-2.4T-A95B
Open Weight Frontier Cloud
67.2
Thinking
78Tool-Use unterstützt, Score: 78
44.2 t/s
$1.89
Report
82
Hermes 4.3 36B
Open Weight Server vLLM
67.1
Standard
67Tool-Use unterstützt, Score: 67
12.7 t/s
Report
83
Codestral 25.08
Restricted Desktop API
65.5
Standard
69Tool-Use unterstützt, Score: 69
192 t/s
$0.05
Report
84
GPT-OSS 20B
Open Weight Desktop vLLM
64.3
Standard
28Tool-Use unterstützt, Score: 28
41.5 t/s
Report
85
GPT-OSS 20B (Thinking)
Open Weight Desktop vLLM
61.6
Thinking
26Tool-Use unterstützt, Score: 26
41.2 t/s
Report
86
Qwen 3 4B
Open Weight Nano llama.cpp
60.9
Adaptive
71Tool-Use unterstützt, Score: 71
74.0 t/s
Report
87
Hermes 3 8B
Restricted Edge llama.cpp
58.8
Standard
56Tool-Use unterstützt, Score: 56
48.0 t/s
Report
88
Qwen 2.5 Coder 7B
Open Weight Edge llama.cpp
55.3
Standard
58Tool-Use unterstützt, Score: 58
37.4 t/s
Report