LLMs head to head 7 use-case scenarios · Equal conditions · one results list
How well do open-weight models hold up against proprietary API models and license-restricted restricted-weight models? Using everyday tasks under identical conditions, all license models and weight classes end up in the same scoreboard.
The strength of the scoreboard lies in comparability. Filtering by size class or model type doesn't hide anything. Non-matching entries are visually de-emphasized but remain in context. This makes it easy to see directly where an open-weight desktop model stands relative to a proprietary frontier model, and whether the performance gap justifies the price gap.
Three license models compete in the scoreboard. Proprietary models neither release training data nor publish weights; access runs exclusively through the vendor's API. Restricted-weight models publish their weights and can be run locally or in the cloud, but the license typically restricts use to academic or non-commercial purposes. Open-weight models are available without restrictions: download, self-host, use commercially.
| # | Report | ||||||
|---|---|---|---|---|---|---|---|
|
1
|
Claude Opus 4.7
Commercial
Frontier
API
|
79.3
|
Adaptive
|
83Tool-Use supported, Score: 83
|
53.1 t/s
|
$2.73
|
Report |
|
2
|
Qwen 3.7 Max
Commercial
Frontier
Cloud
|
78.8
|
Adaptive
|
70Tool-Use supported, Score: 70
|
19.1 t/s
|
$0.64
|
Report |
|
3
|
Kimi K3
Open Weight
Frontier
Cloud
|
78.0
|
Thinking
|
79Tool-Use supported, Score: 79
|
16.1 t/s
|
$3.15
|
Report |
|
4
|
Claude Sonnet 4.6
Commercial
Frontier
API
|
78.0
|
Adaptive
|
66Tool-Use supported, Score: 66
|
43.4 t/s
|
$1.76
|
Report |
|
5
|
MiniMax M3
Open Weight
Frontier
Cloud
|
77.8
|
Adaptive
|
76Tool-Use supported, Score: 76
|
45.3 t/s
|
$0.15
|
Report |
|
6
|
Claude Opus 4.8
Commercial
Frontier
API
|
77.5
|
Thinking
|
71Tool-Use supported, Score: 71
|
63.2 t/s
|
$2.66
|
Report |
|
7
|
Xiaomi MiMo V2.5 Pro
Open Weight
Frontier
Cloud
|
77.0
|
Adaptive
|
68Tool-Use supported, Score: 68
|
27.4 t/s
|
$0.12
|
Report |
|
8
|
Claude Sonnet 5
Commercial
Frontier
API
|
76.9
|
Thinking
|
74Tool-Use supported, Score: 74
|
64.5 t/s
|
$1.41
|
Report |
|
9
|
Kimi K2.7 Code
Open Weight
Frontier
Cloud
|
76.4
|
Thinking
|
71Tool-Use supported, Score: 71
|
70.8 t/s
|
$0.61
|
Report |
|
10
|
GLM-5.1
Open Weight
Frontier
Cloud
|
76.2
|
Adaptive
|
74Tool-Use supported, Score: 74
|
27.0 t/s
|
$0.43
|
Report |
|
11
|
Ornith 1.0 35B (Thinking)
Open Weight
Workstation
LCL
|
76.1
|
Thinking
|
72Tool-Use supported, Score: 72
|
20.1 t/s
|
–
|
Report |
|
12
|
Claude Opus 4.6
Commercial
Frontier
API
|
76.1
|
Adaptive
|
66Tool-Use supported, Score: 66
|
44.0 t/s
|
$2.78
|
Report |
|
13
|
Qwen 3.6 Plus
Commercial
Frontier
Cloud
|
76.0
|
Adaptive
|
59Tool-Use supported, Score: 59
|
18.3 t/s
|
$0.35
|
Report |
|
14
|
Kimi K2.6
Open Weight
Frontier
Cloud
|
75.9
|
Adaptive
|
75Tool-Use supported, Score: 75
|
25.1 t/s
|
$0.75
|
Report |
|
15
|
Qwen 3.6 27B
Open Weight
Workstation
LCL
|
75.8
|
Standard
|
78Tool-Use supported, Score: 78
|
24.7 t/s
|
–
|
Report |
|
16
|
Gemini 2.5 Pro
Commercial
Frontier
API
|
75.3
|
Adaptive
|
74Tool-Use supported, Score: 74
|
30.9 t/s
|
$0.69
|
Report |
|
17
|
MiniMax M2.7
Restricted
Frontier
Cloud
|
75.3
|
Adaptive
|
51Tool-Use supported, Score: 51
|
69.0 t/s
|
$0.11
|
Report |
|
18
|
Claude Haiku 4.5
Commercial
Frontier
API
|
75.1
|
Standard
|
66Tool-Use supported, Score: 66
|
100 t/s
|
$0.45
|
Report |
|
19
|
DeepSeek V4 Pro
Open Weight
Frontier
Cloud
|
75.1
|
Adaptive
|
76Tool-Use supported, Score: 76
|
36.3 t/s
|
$0.09
|
Report |
|
20
|
Kimi K2 Thinking
Open Weight
Frontier
Cloud
|
75.0
|
Thinking
|
76Tool-Use supported, Score: 76
|
34.9 t/s
|
$0.37
|
Report |
|
21
|
Gemma 4 26B-A4B Instruct (Thinking)
Open Weight
Workstation
LCL
|
74.9
|
Thinking
|
74Tool-Use supported, Score: 74
|
27.3 t/s
|
–
|
Report |
|
22
|
Claude Sonnet 4.5
Commercial
Frontier
API
|
74.9
|
Adaptive
|
72Tool-Use supported, Score: 72
|
50.0 t/s
|
$1.24
|
Report |
|
23
|
GLM-5
Open Weight
Frontier
Cloud
|
74.8
|
Adaptive
|
74Tool-Use supported, Score: 74
|
39.8 t/s
|
$0.25
|
Report |
|
24
|
Gemma 4 31B Instruct (Thinking)
Open Weight
Workstation
LCL
|
74.8
|
Thinking
|
73Tool-Use supported, Score: 73
|
6.6 t/s
|
–
|
Report |
|
25
|
DeepSeek V4 Flash
Open Weight
Frontier
Cloud
|
74.6
|
Adaptive
|
78Tool-Use supported, Score: 78
|
44.0 t/s
|
$0.02
|
Report |
|
26
|
GPT-5
Commercial
Frontier
API
|
74.5
|
Adaptive
|
62Tool-Use supported, Score: 62
|
34.9 t/s
|
$1.75
|
Report |
|
27
|
GLM-4.7
Restricted
Frontier
Cloud
|
74.5
|
Adaptive
|
63Tool-Use supported, Score: 63
|
26.1 t/s
|
$0.25
|
Report |
|
28
|
Qwen 3.6 35B-A3B
Open Weight
Workstation
LCL
|
74.4
|
Standard
|
74Tool-Use supported, Score: 74
|
87.1 t/s
|
–
|
Report |
|
29
|
GLM-5.2
Open Weight
Frontier
Cloud
|
74.3
|
Adaptive
|
76Tool-Use supported, Score: 76
|
37.6 t/s
|
$0.48
|
Report |
|
30
|
GPT-5.5
Commercial
Frontier
API
|
74.3
|
Thinking
|
71Tool-Use supported, Score: 71
|
38.6 t/s
|
$2.99
|
Report |
|
31
|
Qwen 3.5 35B-A3B (Unsloth)
Open Weight
Workstation
LCL
|
74.3
|
Adaptive
|
72Tool-Use supported, Score: 72
|
69.3 t/s
|
–
|
Report |
|
32
|
Qwen 3.6 35B-A3B (Thinking)
Open Weight
Workstation
LCL
|
74.2
|
Thinking
|
72Tool-Use supported, Score: 72
|
41.2 t/s
|
–
|
Report |
|
33
|
Qwen 3 Coder Next
Open Weight
Workstation
LCL
|
74.1
|
Standard
|
71Tool-Use supported, Score: 71
|
48.5 t/s
|
–
|
Report |
|
34
|
Mistral Medium 3.5
Open Weight
Frontier
API
|
73.9
|
Standard
|
74Tool-Use supported, Score: 74
|
135 t/s
|
$0.48
|
Report |
|
35
|
GPT-OSS 120B
Open Weight
Server
LCL
|
73.8
|
Standard
|
66Tool-Use supported, Score: 66
|
27.5 t/s
|
–
|
Report |
|
36
|
Gemini 3.5 Flash
Commercial
Frontier
API
|
73.7
|
Adaptive
|
72Tool-Use supported, Score: 72
|
51.9 t/s
|
$0.47
|
Report |
|
37
|
Qwen 3.6 27B (Thinking)
Open Weight
Workstation
LCL
|
73.5
|
Thinking
|
62Tool-Use supported, Score: 62
|
10.3 t/s
|
–
|
Report |
|
38
|
GPT-OSS 120B (Thinking)
Open Weight
Server
LCL
|
73.4
|
Thinking
|
72Tool-Use supported, Score: 72
|
26.9 t/s
|
–
|
Report |
|
39
|
Gemma 4 ARA 26B-A4B (ARA-Abliterated)
Open Weight
Workstation
LCL
|
73.3
|
Standard
|
76Tool-Use supported, Score: 76
|
53.5 t/s
|
–
|
Report |
|
40
|
Kimi K2.5
Open Weight
Frontier
Cloud
|
73.3
|
Thinking
|
77Tool-Use supported, Score: 77
|
21.1 t/s
|
$0.33
|
Report |
|
41
|
Gemma 4 26B-A4B Instruct
Open Weight
Workstation
LCL
|
73.3
|
Standard
|
73Tool-Use supported, Score: 73
|
70.7 t/s
|
–
|
Report |
|
42
|
Gemma 4 31B Instruct
Open Weight
Workstation
Cloud
|
73.2
|
Standard
|
74Tool-Use supported, Score: 74
|
31.2 t/s
|
$0.02
|
Report |
|
43
|
Qwen 3.5 397B A17B
Open Weight
Frontier
Cloud
|
73.2
|
Adaptive
|
74Tool-Use supported, Score: 74
|
20.4 t/s
|
$0.46
|
Report |
|
44
|
Mistral 3 Large
Open Weight
Frontier
API
|
73.2
|
Standard
|
60Tool-Use supported, Score: 60
|
55.1 t/s
|
$0.43
|
Report |
|
45
|
GLM 4.6
Restricted
Frontier
Cloud
|
73.1
|
Standard
|
76Tool-Use supported, Score: 76
|
17.8 t/s
|
$0.27
|
Report |
|
46
|
GPT-5.4
Commercial
Frontier
API
|
72.8
|
Standard
|
57Tool-Use supported, Score: 57
|
78.7 t/s
|
$0.98
|
Report |
|
47
|
GLM-5 Turbo
Commercial
Frontier
Cloud
|
72.7
|
Adaptive
|
79Tool-Use supported, Score: 79
|
20.9 t/s
|
$0.45
|
Report |
|
48
|
Gemma 4 31B Ortenzya Creative Wordsmith (Thinking)
Open Weight
Workstation
LCL
|
72.6
|
Thinking
|
72Tool-Use supported, Score: 72
|
6.7 t/s
|
–
|
Report |
|
49
|
Gemma 4 31B Ortenzya Creative Wordsmith
Open Weight
Workstation
LCL
|
72.5
|
Standard
|
73Tool-Use supported, Score: 73
|
16.4 t/s
|
–
|
Report |
|
50
|
NVIDIA Nemotron 3 Ultra 550B A55B
Open Weight
Frontier
Cloud
|
72.5
|
Adaptive
|
55Tool-Use supported, Score: 55
|
98.9 t/s
|
$0.26
|
Report |
|
51
|
Qwen 3.6 35B-A3B (Unsloth)
Open Weight
Desktop
LCL
|
72.5
|
Adaptive
|
66Tool-Use supported, Score: 66
|
66.4 t/s
|
–
|
Report |
|
52
|
GPT-5 Mini
Commercial
Nano
API
|
72.3
|
Standard
|
17Tool-Use supported, Score: 17
|
34.6 t/s
|
$0.23
|
Report |
|
53
|
Ornith 1.0 35B
Open Weight
Workstation
LCL
|
72.1
|
Standard
|
52Tool-Use supported, Score: 52
|
40.1 t/s
|
–
|
Report |
|
54
|
Grok 4.5
Commercial
Frontier
API
|
72.1
|
Standard
|
74Tool-Use supported, Score: 74
|
46.7 t/s
|
$0.42
|
Report |
|
55
|
Gemma 4 31B Instruct
Open Weight
Workstation
LCL
|
72.0
|
Standard
|
68Tool-Use supported, Score: 68
|
14.1 t/s
|
–
|
Report |
|
56
|
DeepSeek V3.2
Open Weight
Frontier
Cloud
|
71.9
|
Standard
|
67Tool-Use supported, Score: 67
|
46.3 t/s
|
$0.02
|
Report |
|
57
|
Gemma 4 E4B
Open Weight
Edge
LCL
|
71.9
|
Adaptive
|
73Tool-Use supported, Score: 73
|
48.2 t/s
|
–
|
Report |
|
58
|
Qwen 3.6 35B-A3B (Uncensored)
Open Weight
Desktop
LCL
|
71.7
|
Adaptive
|
60Tool-Use supported, Score: 60
|
53.9 t/s
|
–
|
Report |
|
59
|
Kimi K2
Open Weight
Frontier
Cloud
|
71.3
|
Standard
|
71Tool-Use supported, Score: 71
|
21.3 t/s
|
$0.12
|
Report |
|
60
|
Gemma 4 12B Instruct (Unsloth)
Open Weight
Desktop
LCL
|
70.7
|
Adaptive
|
62Tool-Use supported, Score: 62
|
13.1 t/s
|
–
|
Report |
|
61
|
Mistral Small 4
Open Weight
Workstation
API
|
70.7
|
Standard
|
53Tool-Use supported, Score: 53
|
177 t/s
|
$0.02
|
Report |
|
62
|
Xiaomi MiMo V2.5
Open Weight
Frontier
Cloud
|
70.6
|
Adaptive
|
79Tool-Use supported, Score: 79
|
37.5 t/s
|
$0.23
|
Report |
|
63
|
Hermes 4 70B
Open Weight
Server
Cloud
|
70.4
|
Adaptive
|
66Tool-Use supported, Score: 66
|
75.8 t/s
|
$0.02
|
Report |
|
64
|
Llama 3.3 Nemotron Super 49B v1.5
Open Weight
Server
Cloud
|
70.2
|
Adaptive
|
73Tool-Use supported, Score: 73
|
20.7 t/s
|
$0.04
|
Report |
|
65
|
Devstral 2
Open Weight
Frontier
API
|
70.2
|
Standard
|
55Tool-Use supported, Score: 55
|
56.2 t/s
|
$0.12
|
Report |
|
66
|
GPT-5.4 Mini
Commercial
Frontier
API
|
70.2
|
Standard
|
68Tool-Use supported, Score: 68
|
120 t/s
|
$0.25
|
Report |
|
67
|
Grok 4 (Non-Reasoning)
Commercial
Frontier
API
|
70.1
|
Adaptive
|
55Tool-Use supported, Score: 55
|
182 t/s
|
$0.17
|
Report |
|
68
|
Grok 4.20 (Reasoning)
Commercial
Frontier
API
|
70.1
|
Thinking
|
75Tool-Use supported, Score: 75
|
41.0 t/s
|
$0.15
|
Report |
|
69
|
Qwen 3.5 4B (Unsloth)
Open Weight
Nano
LCL
|
69.9
|
Adaptive
|
78Tool-Use supported, Score: 78
|
51.3 t/s
|
–
|
Report |
|
70
|
o4-mini
Commercial
Frontier
API
|
69.5
|
Thinking
|
63Tool-Use supported, Score: 63
|
69.7 t/s
|
$0.40
|
Report |
|
71
|
Qwen 3 32B
Open Weight
Workstation
Cloud
|
69.3
|
Adaptive
|
65Tool-Use supported, Score: 65
|
161 t/s
|
$0.06
|
Report |
|
72
|
Laguna S 2.1
Open Weight
Server
LCL
|
69.1
|
Thinking
|
63Tool-Use supported, Score: 63
|
17.2 t/s
|
$0.02
|
Report |
|
73
|
Gemma 3 12B IT
Restricted
Desktop
LCL
|
68.9
|
Standard
|
64Tool-Use supported, Score: 64
|
39.1 t/s
|
–
|
Report |
|
74
|
GPT-5.4 Nano
Commercial
Frontier
API
|
68.9
|
Standard
|
58Tool-Use supported, Score: 58
|
123 t/s
|
$0.07
|
Report |
|
75
|
Qwen 3.5 9B (Unsloth)
Open Weight
Edge
LCL
|
68.7
|
Adaptive
|
57Tool-Use supported, Score: 57
|
34.3 t/s
|
–
|
Report |
|
76
|
NVIDIA Nemotron 3 Nano 30B A3B
Open Weight
Workstation
Cloud
|
68.3
|
Adaptive
|
66Tool-Use supported, Score: 66
|
36.1 t/s
|
$0.02
|
Report |
|
77
|
Hermes 4 14B (Abliterated)
Open Weight
Desktop
LCL
|
68.2
|
Adaptive
|
70Tool-Use supported, Score: 70
|
25.4 t/s
|
–
|
Report |
|
78
|
o3-mini
Commercial
Frontier
API
|
68.0
|
Thinking
|
73Tool-Use supported, Score: 73
|
71.2 t/s
|
$0.37
|
Report |
|
79
|
Hermes 4.3 36B (Thinking)
Open Weight
Server
LCL
|
68.0
|
Thinking
|
67Tool-Use supported, Score: 67
|
13.3 t/s
|
–
|
Report |
|
80
|
GPT-4o
Commercial
Frontier
API
|
68.0
|
Standard
|
67Tool-Use supported, Score: 67
|
150 t/s
|
$0.48
|
Report |
|
81
|
Hermes 4 405B
Restricted
Frontier
Cloud
|
67.8
|
Adaptive
|
74Tool-Use supported, Score: 74
|
39.5 t/s
|
$0.15
|
Report |
|
82
|
Qwen 3 14B
Open Weight
Desktop
LCL
|
67.7
|
Adaptive
|
66Tool-Use supported, Score: 66
|
23.9 t/s
|
–
|
Report |
|
83
|
OpenAI o1
Commercial
Frontier
API
|
67.5
|
Thinking
|
77Tool-Use supported, Score: 77
|
52.5 t/s
|
$6.46
|
Report |
|
84
|
Hermes 4 14B
Open Weight
Desktop
LCL
|
67.3
|
Adaptive
|
67Tool-Use supported, Score: 67
|
30.3 t/s
|
–
|
Report |
|
85
|
Grok 4.3
Commercial
Frontier
API
|
67.3
|
Thinking
|
62Tool-Use supported, Score: 62
|
65.7 t/s
|
$0.12
|
Report |
|
86
|
Hermes 4.3 36B
Open Weight
Server
LCL
|
67.1
|
Standard
|
67Tool-Use supported, Score: 67
|
12.7 t/s
|
–
|
Report |
|
87
|
Codestral 25.08
Restricted
Frontier
API
|
65.5
|
Standard
|
69Tool-Use supported, Score: 69
|
192 t/s
|
$0.03
|
Report |
|
88
|
GPT-4o Mini
Commercial
Frontier
API
|
65.0
|
Standard
|
66Tool-Use supported, Score: 66
|
80.8 t/s
|
$0.03
|
Report |
|
89
|
Llama 3.3 70B Versatile
Restricted
Server
Cloud
|
64.3
|
Standard
|
42Tool-Use supported, Score: 42
|
276 t/s
|
$0.03
|
Report |
|
90
|
Command A+
Open Weight
Frontier
Cloud
|
61.9
|
Thinking
|
6Tool-Use supported, Score: 6
|
76.5 t/s
|
–
|
Report |
|
91
|
Qwen 3 4B
Open Weight
Nano
LCL
|
60.9
|
Adaptive
|
71Tool-Use supported, Score: 71
|
74.0 t/s
|
–
|
Report |
|
92
|
Hermes 3 8B
Restricted
Edge
LCL
|
58.8
|
Standard
|
56Tool-Use supported, Score: 56
|
48.0 t/s
|
–
|
Report |
|
93
|
GPT-OSS 20B
Open Weight
Desktop
Cloud
|
56.2
|
Adaptive
|
0Tool-Use supported, Score: 0
|
406 t/s
|
$0.03
|
Report |
|
94
|
Qwen 2.5 Coder 7B
Open Weight
Nano
LCL
|
56.1
|
Standard
|
58Tool-Use supported, Score: 58
|
50.1 t/s
|
–
|
Report |
|
95
|
Llama 4 Scout 17B
Restricted
Server
Cloud
|
55.2
|
Standard
|
33Tool-Use supported, Score: 33
|
355 t/s
|
$0.01
|
Report |