LLMs im direkten Vergleich 7 Einsatzszenarien · Gleiche Bedingungen · eine Ergebnisliste

Wie gut schlagen sich freie Open-Weight-Modelle gegen proprietäre API-Modelle und lizenzlich eingeschränkte Restricted-Weight-Modelle? Anhand alltäglicher Aufgaben und unter denselben Bedingungen landen alle Lizenzmodelle und Gewichtsklassen im selben Scoreboard.

Scoreboard

94 Modelle im Scoreboard

# Report
1
Claude Opus 4.7
Commercial Frontier API
79.3
Adaptive
83Tool-Use unterstützt, Score: 83
53.1 t/s
$2.73
Report
2
Qwen 3.7 Max
Commercial Frontier LCL
78.8
Adaptive
70Tool-Use unterstützt, Score: 70
19.1 t/s
$0.64
Report
3
Kimi K3
Open Weight Frontier LCL
78.0
Thinking
79Tool-Use unterstützt, Score: 79
16.1 t/s
$3.15
Report
4
Claude Sonnet 4.6
Commercial Frontier API
78.0
Adaptive
66Tool-Use unterstützt, Score: 66
43.4 t/s
$1.76
Report
5
MiniMax M3
Open Weight Frontier LCL
77.8
Adaptive
76Tool-Use unterstützt, Score: 76
45.3 t/s
$0.15
Report
6
Claude Opus 4.8
Commercial Frontier API
77.5
Thinking
71Tool-Use unterstützt, Score: 71
63.2 t/s
$2.66
Report
7
Xiaomi MiMo V2.5 Pro
Open Weight Frontier LCL
77.0
Adaptive
68Tool-Use unterstützt, Score: 68
27.4 t/s
$0.12
Report
8
Claude Sonnet 5
Commercial Frontier API
76.9
Thinking
74Tool-Use unterstützt, Score: 74
64.5 t/s
$1.41
Report
9
Ornith 1.0 35B (FP8) (Thinking)
Open Weight Workstation LCL
76.9
Thinking
74Tool-Use unterstützt, Score: 74
21.6 t/s
Report
10
Kimi K2.7 Code
Open Weight Frontier LCL
76.4
Thinking
71Tool-Use unterstützt, Score: 71
70.8 t/s
$0.61
Report
11
GLM-5.1
Open Weight Frontier LCL
76.2
Adaptive
74Tool-Use unterstützt, Score: 74
27.0 t/s
$0.43
Report
12
Claude Opus 4.6
Commercial Frontier API
76.1
Adaptive
66Tool-Use unterstützt, Score: 66
44.0 t/s
$2.78
Report
13
Qwen 3.6 Plus
Commercial Frontier LCL
76.0
Adaptive
59Tool-Use unterstützt, Score: 59
18.3 t/s
$0.35
Report
14
Kimi K2.6
Open Weight Frontier LCL
75.9
Adaptive
75Tool-Use unterstützt, Score: 75
25.1 t/s
$0.75
Report
15
Gemini 2.5 Pro
Commercial Frontier API
75.3
Adaptive
74Tool-Use unterstützt, Score: 74
30.9 t/s
$0.69
Report
16
MiniMax M2.7
Restricted Frontier LCL
75.3
Adaptive
51Tool-Use unterstützt, Score: 51
69.0 t/s
$0.11
Report
17
Gemma 4 31B Instruct
Open Weight Workstation LCL
75.2
Standard
69Tool-Use unterstützt, Score: 69
16.9 t/s
Report
18
Claude Haiku 4.5
Commercial Frontier API
75.1
Standard
66Tool-Use unterstützt, Score: 66
100 t/s
$0.45
Report
19
DeepSeek V4 Pro
Open Weight Frontier LCL
75.1
Adaptive
76Tool-Use unterstützt, Score: 76
36.3 t/s
$0.09
Report
20
Kimi K2 Thinking
Open Weight Frontier LCL
75.0
Thinking
76Tool-Use unterstützt, Score: 76
34.9 t/s
$0.37
Report
21
Claude Sonnet 4.5
Commercial Frontier API
74.9
Adaptive
72Tool-Use unterstützt, Score: 72
50.0 t/s
$1.24
Report
22
GLM-5
Open Weight Frontier LCL
74.8
Adaptive
74Tool-Use unterstützt, Score: 74
39.8 t/s
$0.25
Report
23
DeepSeek V4 Flash
Open Weight Frontier LCL
74.6
Adaptive
78Tool-Use unterstützt, Score: 78
44.0 t/s
$0.02
Report
24
GPT-5
Commercial Frontier API
74.5
Adaptive
62Tool-Use unterstützt, Score: 62
34.9 t/s
$1.75
Report
25
GLM-4.7
Restricted Frontier LCL
74.5
Adaptive
63Tool-Use unterstützt, Score: 63
26.1 t/s
$0.25
Report
26
GLM-5.2
Open Weight Frontier LCL
74.3
Adaptive
76Tool-Use unterstützt, Score: 76
37.6 t/s
$0.48
Report
27
GPT-5.5
Commercial Frontier API
74.3
Thinking
71Tool-Use unterstützt, Score: 71
38.6 t/s
$2.99
Report
28
Qwen 3.5 35B-A3B (Unsloth)
Open Weight Workstation LCL
74.3
Adaptive
72Tool-Use unterstützt, Score: 72
69.3 t/s
Report
29
Qwen 3 Coder Next
Open Weight Workstation LCL
74.1
Standard
71Tool-Use unterstützt, Score: 71
48.5 t/s
Report
30
GPT-OSS 120B
Open Weight Workstation LCL
74.1
Adaptive
69Tool-Use unterstützt, Score: 69
369 t/s
$0.05
Report
31
Gemma 4 26B-A4B Instruct
Open Weight Workstation LCL
74.0
Standard
72Tool-Use unterstützt, Score: 72
47.8 t/s
Report
32
Qwen3.6 27B Instruct (Thinking)
Open Weight Workstation LCL
73.9
Thinking
20Tool-Use unterstützt, Score: 20
7.2 t/s
Report
33
Mistral Medium 3.5
Open Weight Frontier API
73.9
Standard
74Tool-Use unterstützt, Score: 74
135 t/s
$0.48
Report
34
Gemma 4 Ortenzya Creative Wordsmith 31B
Open Weight Workstation LCL
73.8
Standard
75Tool-Use unterstützt, Score: 75
17.5 t/s
Report
35
Gemini 3.5 Flash
Commercial Frontier API
73.7
Adaptive
72Tool-Use unterstützt, Score: 72
51.9 t/s
$0.47
Report
36
Ornith 1.0 35B (FP8)
Open Weight Workstation LCL
73.4
Standard
71Tool-Use unterstützt, Score: 71
49.3 t/s
Report
37
Qwen 3.6 27B NVFP4 (vLLM, Dense, MTP)
Open Weight Workstation LCL
73.4
Standard
65Tool-Use unterstützt, Score: 65
22.6 t/s
Report
38
Gemma 4 ARA 26B-A4B (ARA-Abliterated)
Open Weight Workstation LCL
73.3
Standard
76Tool-Use unterstützt, Score: 76
53.5 t/s
Report
39
Kimi K2.5
Open Weight Frontier LCL
73.3
Thinking
77Tool-Use unterstützt, Score: 77
21.1 t/s
$0.33
Report
40
Gemma 4 31B Instruct
Open Weight Workstation LCL
73.2
Standard
74Tool-Use unterstützt, Score: 74
31.2 t/s
$0.02
Report
41
Qwen 3.6 35B-A3B NVFP4 (vLLM, MoE, MTP) (Thinking)
Open Weight Workstation LCL
73.2
Thinking
68Tool-Use unterstützt, Score: 68
36.6 t/s
Report
42
Qwen 3.5 397B A17B
Open Weight Frontier LCL
73.2
Adaptive
74Tool-Use unterstützt, Score: 74
20.4 t/s
$0.46
Report
43
Mistral 3 Large
Open Weight Frontier API
73.2
Standard
60Tool-Use unterstützt, Score: 60
55.1 t/s
$0.43
Report
44
GLM 4.6
Restricted Frontier LCL
73.1
Standard
76Tool-Use unterstützt, Score: 76
17.8 t/s
$0.27
Report
45
GPT-5.4
Commercial Frontier API
72.8
Standard
57Tool-Use unterstützt, Score: 57
78.7 t/s
$0.98
Report
46
GLM-5 Turbo
Commercial Frontier LCL
72.7
Adaptive
79Tool-Use unterstützt, Score: 79
20.9 t/s
$0.45
Report
47
NVIDIA Nemotron 3 Ultra 550B A55B
Open Weight Frontier LCL
72.5
Adaptive
55Tool-Use unterstützt, Score: 55
98.9 t/s
$0.26
Report
48
Qwen 3.6 35B-A3B (Unsloth)
Open Weight Desktop LCL
72.5
Adaptive
66Tool-Use unterstützt, Score: 66
66.4 t/s
Report
49
GPT-5 Mini
Commercial Nano API
72.3
Standard
17Tool-Use unterstützt, Score: 17
34.6 t/s
$0.23
Report
50
Grok 4.5
Commercial Frontier API
72.1
Standard
74Tool-Use unterstützt, Score: 74
46.7 t/s
$0.42
Report
51
DeepSeek V3.2
Open Weight Frontier LCL
71.9
Standard
67Tool-Use unterstützt, Score: 67
46.3 t/s
$0.02
Report
52
Gemma 4 Ortenzya Creative Wordsmith 31B (Thinking)
Open Weight Workstation LCL
71.9
Thinking
66Tool-Use unterstützt, Score: 66
7.1 t/s
Report
53
Gemma 4 E4B
Open Weight Edge LCL
71.9
Adaptive
73Tool-Use unterstützt, Score: 73
48.2 t/s
Report
54
Qwen 3.6 27B NVFP4 (vLLM, Dense, MTP) (Thinking)
Open Weight Workstation LCL
71.7
Thinking
60Tool-Use unterstützt, Score: 60
8.5 t/s
Report
55
Qwen 3.6 35B-A3B (Uncensored)
Open Weight Desktop LCL
71.7
Adaptive
60Tool-Use unterstützt, Score: 60
53.9 t/s
Report
56
Gemma 4 31B Instruct (Thinking)
Open Weight Workstation LCL
71.6
Thinking
69Tool-Use unterstützt, Score: 69
7.5 t/s
Report
57
Kimi K2
Open Weight Frontier LCL
71.3
Standard
71Tool-Use unterstützt, Score: 71
21.3 t/s
$0.12
Report
58
Qwen3.6 27B Instruct
Open Weight Workstation LCL
71.1
Standard
64Tool-Use unterstützt, Score: 64
22.1 t/s
Report
59
Gemma 4 12B Instruct (Unsloth)
Open Weight Desktop LCL
70.7
Adaptive
62Tool-Use unterstützt, Score: 62
13.1 t/s
Report
60
Mistral Small 4
Open Weight Workstation API
70.7
Standard
53Tool-Use unterstützt, Score: 53
177 t/s
$0.02
Report
61
Xiaomi MiMo V2.5
Open Weight Frontier LCL
70.6
Adaptive
79Tool-Use unterstützt, Score: 79
37.5 t/s
$0.23
Report
62
Hermes 4 70B
Open Weight Server LCL
70.4
Adaptive
66Tool-Use unterstützt, Score: 66
75.8 t/s
$0.02
Report
63
Qwen 3.6 35B-A3B NVFP4 (vLLM, MoE, MTP)
Open Weight Workstation LCL
70.3
Standard
63Tool-Use unterstützt, Score: 63
79.1 t/s
Report
64
Gemma 4 26B-A4B Instruct (Thinking)
Open Weight Workstation LCL
70.3
Thinking
69Tool-Use unterstützt, Score: 69
18.6 t/s
Report
65
Llama 3.3 Nemotron Super 49B v1.5
Open Weight Server LCL
70.2
Adaptive
73Tool-Use unterstützt, Score: 73
20.7 t/s
$0.04
Report
66
Gemma 4 31B Creative Wordsmith (Uncensored)
Open Weight Workstation LCL
70.2
Adaptive
66Tool-Use unterstützt, Score: 66
9.4 t/s
Report
67
Devstral 2
Open Weight Frontier API
70.2
Standard
55Tool-Use unterstützt, Score: 55
56.2 t/s
$0.12
Report
68
GPT-5.4 Mini
Commercial Frontier API
70.2
Standard
68Tool-Use unterstützt, Score: 68
120 t/s
$0.25
Report
69
Grok 4 (Non-Reasoning)
Commercial Frontier API
70.1
Adaptive
55Tool-Use unterstützt, Score: 55
182 t/s
$0.17
Report
70
Grok 4.20 (Reasoning)
Commercial Frontier API
70.1
Thinking
75Tool-Use unterstützt, Score: 75
41.0 t/s
$0.15
Report
71
Qwen 3.5 4B (Unsloth)
Open Weight Nano LCL
69.9
Adaptive
78Tool-Use unterstützt, Score: 78
51.3 t/s
Report
72
o4-mini
Commercial Frontier API
69.5
Thinking
63Tool-Use unterstützt, Score: 63
69.7 t/s
$0.40
Report
73
Qwen 3 32B
Open Weight Workstation LCL
69.3
Adaptive
65Tool-Use unterstützt, Score: 65
161 t/s
$0.06
Report
74
Gemma 3 12B IT
Restricted Desktop LCL
68.9
Standard
64Tool-Use unterstützt, Score: 64
39.1 t/s
Report
75
GPT-5.4 Nano
Commercial Frontier API
68.9
Standard
58Tool-Use unterstützt, Score: 58
123 t/s
$0.07
Report
76
Qwen 3.5 9B (Unsloth)
Open Weight Edge LCL
68.7
Adaptive
57Tool-Use unterstützt, Score: 57
34.3 t/s
Report
77
NVIDIA Nemotron 3 Nano 30B A3B
Open Weight Workstation LCL
68.3
Adaptive
66Tool-Use unterstützt, Score: 66
36.1 t/s
$0.02
Report
78
Hermes 4 14B (Abliterated)
Open Weight Desktop LCL
68.2
Adaptive
70Tool-Use unterstützt, Score: 70
25.4 t/s
Report
79
o3-mini
Commercial Frontier API
68.0
Thinking
73Tool-Use unterstützt, Score: 73
71.2 t/s
$0.37
Report
80
GPT-4o
Commercial Frontier API
68.0
Standard
67Tool-Use unterstützt, Score: 67
150 t/s
$0.48
Report
81
Hermes 4 405B
Restricted Frontier LCL
67.8
Adaptive
74Tool-Use unterstützt, Score: 74
39.5 t/s
$0.15
Report
82
Qwen 3 14B
Open Weight Desktop LCL
67.7
Adaptive
66Tool-Use unterstützt, Score: 66
23.9 t/s
Report
83
OpenAI o1
Commercial Frontier API
67.5
Thinking
77Tool-Use unterstützt, Score: 77
52.5 t/s
$6.46
Report
84
Hermes 4 14B
Open Weight Desktop LCL
67.3
Adaptive
67Tool-Use unterstützt, Score: 67
30.3 t/s
Report
85
Grok 4.3
Commercial Frontier API
67.3
Thinking
62Tool-Use unterstützt, Score: 62
65.7 t/s
$0.12
Report
86
Codestral 25.08
Restricted Frontier API
65.5
Standard
69Tool-Use unterstützt, Score: 69
192 t/s
$0.03
Report
87
GPT-4o Mini
Commercial Frontier API
65.0
Standard
66Tool-Use unterstützt, Score: 66
80.8 t/s
$0.03
Report
88
Llama 3.3 70B Versatile
Restricted Server LCL
64.3
Standard
42Tool-Use unterstützt, Score: 42
276 t/s
$0.03
Report
89
Command A+
Open Weight Frontier LCL
61.9
Thinking
6Tool-Use unterstützt, Score: 6
76.5 t/s
Report
90
Qwen 3 4B
Open Weight Nano LCL
60.9
Adaptive
71Tool-Use unterstützt, Score: 71
74.0 t/s
Report
91
Hermes 3 8B
Restricted Edge LCL
58.8
Standard
56Tool-Use unterstützt, Score: 56
48.0 t/s
Report
92
GPT-OSS 20B
Open Weight Desktop LCL
56.2
Adaptive
0Tool-Use unterstützt, Score: 0
406 t/s
$0.03
Report
93
Qwen 2.5 Coder 7B
Open Weight Nano LCL
56.1
Standard
58Tool-Use unterstützt, Score: 58
50.1 t/s
Report
94
Llama 4 Scout 17B
Restricted Server LCL
55.2
Standard
33Tool-Use unterstützt, Score: 33
355 t/s
$0.01
Report