LLMs head to head 7 use-case scenarios · Equal conditions · one results list
How well do open-weight models hold up against proprietary API models and license-restricted restricted-weight models? Using everyday tasks under identical conditions, all license models and weight classes land in the same scoreboard.
The total score shows the best all-rounder. The profiles below answer the real everyday question: is this model sufficient for your task — even if it's a small, local specialist rather than a cloud giant?
The scoreboard's strength lies in comparability. Filtering by size class or model type doesn't hide anything. Non-matching entries are visually de-emphasized but remain in context. This makes it easy to see directly where an open-weight desktop model stands relative to a proprietary frontier model, and whether the performance difference justifies the price difference.
Three license models compete in the scoreboard. Proprietary models neither release training data nor publish weights; access runs exclusively through the vendor's API. Restricted-weight models publish their weights and can be run locally or in the cloud, but the license typically restricts use to academic or non-commercial purposes. Open-weight models are available without restrictions: download, self-host, use commercially.
| # | Report | ||||||
|---|---|---|---|---|---|---|---|
|
1
|
GLM-5.3-Flash
Open Weight
Server
Cloud
|
80.6
|
Thinking
|
78Tool-Use supported, Score: 78
|
52.7 t/s
|
$0.11
|
Report |
|
2
|
MiniMax M3
Open Weight
Server
Cloud
|
80.2
|
Adaptive
|
81Tool-Use supported, Score: 81
|
108 t/s
|
$0.16
|
Report |
|
3
|
Qwen 3.8 27B (Thinking)
Open Weight
Workstation
vLLM
|
80.0
|
Thinking
|
87Tool-Use supported, Score: 87
|
21.6 t/s
|
–
|
Report |
|
4
|
Claude Sonnet 4.6
Commercial
Frontier
API
|
79.8
|
Adaptive
|
75Tool-Use supported, Score: 75
|
51.1 t/s
|
$1.77
|
Report |
|
5
|
Qwen3.8-Flash
Open Weight
Server
Cloud
|
79.5
|
Thinking
|
76Tool-Use supported, Score: 76
|
50.9 t/s
|
$0.05
|
Report |
|
6
|
Claude Sonnet 5
Commercial
Frontier
API
|
79.0
|
Thinking
|
80Tool-Use supported, Score: 80
|
82.5 t/s
|
$1.43
|
Report |
|
7
|
Xiaomi MiMo V2.5 Pro
Open Weight
Frontier
Cloud
|
79.0
|
Adaptive
|
74Tool-Use supported, Score: 74
|
48.8 t/s
|
$0.12
|
Report |
|
8
|
Kimi K3
Open Weight
Frontier
Cloud
|
78.8
|
Thinking
|
81Tool-Use supported, Score: 81
|
48.3 t/s
|
$3.30
|
Report |
|
9
|
Claude Opus 5
Commercial
Frontier
API
|
78.8
|
Thinking
|
80Tool-Use supported, Score: 80
|
63.8 t/s
|
$4.62
|
Report |
|
10
|
Claude Opus 4.8
Commercial
Frontier
API
|
77.9
|
Thinking
|
72Tool-Use supported, Score: 72
|
63.3 t/s
|
$2.71
|
Report |
|
11
|
Swift Qwen 3.8 27B
Open Weight
Workstation
vLLM
|
77.8
|
Thinking
|
74Tool-Use supported, Score: 74
|
15.5 t/s
|
–
|
Report |
|
12
|
Gemini 3.5 Flash
Commercial
Frontier
API
|
77.2
|
Adaptive
|
76Tool-Use supported, Score: 76
|
58.7 t/s
|
$0.49
|
Report |
|
13
|
Meta Muse Spark 1.2
Commercial
Frontier
Cloud
|
76.7
|
Adaptive
|
70Tool-Use supported, Score: 70
|
115 t/s
|
$0.54
|
Report |
|
14
|
Kimi K2.5
Open Weight
Frontier
Cloud
|
76.6
|
Thinking
|
77Tool-Use supported, Score: 77
|
37.7 t/s
|
$0.39
|
Report |
|
15
|
GPT-5.5
Commercial
Frontier
API
|
76.5
|
Thinking
|
79Tool-Use supported, Score: 79
|
59.3 t/s
|
$2.97
|
Report |
|
16
|
GLM 4.6
Restricted
Server
Cloud
|
76.2
|
Standard
|
53Tool-Use supported, Score: 53
|
46.1 t/s
|
$0.25
|
Report |
|
17
|
Qwen 3.8 27B Uncensored (Thinking)
Open Weight
Workstation
vLLM
|
76.1
|
Thinking
|
74Tool-Use supported, Score: 74
|
16.1 t/s
|
–
|
Report |
|
18
|
Ornith 1.5 35B-A3B
Open Weight
Workstation
vLLM
|
76.0
|
Thinking
|
76Tool-Use supported, Score: 76
|
61.0 t/s
|
–
|
Report |
|
19
|
Qwen 3.7 Max
Commercial
Frontier
Cloud
|
75.8
|
Adaptive
|
60Tool-Use supported, Score: 60
|
62.6 t/s
|
$0.75
|
Report |
|
20
|
Ornith 1.0 35B
Open Weight
Workstation
vLLM
|
75.8
|
Thinking
|
63Tool-Use supported, Score: 63
|
35.5 t/s
|
–
|
Report |
|
21
|
Qwen 3.6 27B (Thinking)
Open Weight
Workstation
vLLM
|
75.6
|
Thinking
|
73Tool-Use supported, Score: 73
|
25.2 t/s
|
–
|
Report |
|
22
|
Kimi K2.6
Open Weight
Frontier
Cloud
|
75.6
|
Adaptive
|
74Tool-Use supported, Score: 74
|
49.4 t/s
|
$0.96
|
Report |
|
23
|
GLM-5.2
Open Weight
Server
Cloud
|
75.5
|
Adaptive
|
51Tool-Use supported, Score: 51
|
65.7 t/s
|
$0.40
|
Report |
|
24
|
Qwen 3.6 Plus
Commercial
Frontier
Cloud
|
75.5
|
Adaptive
|
68Tool-Use supported, Score: 68
|
56.3 t/s
|
$0.41
|
Report |
|
25
|
Upstage Solar Pro4
Commercial
Frontier
Cloud
|
75.4
|
Standard
|
71Tool-Use supported, Score: 71
|
35.9 t/s
|
$0.01
|
Report |
|
26
|
Qwen 3.8 27B
Open Weight
Workstation
vLLM
|
75.2
|
Standard
|
70Tool-Use supported, Score: 70
|
17.4 t/s
|
–
|
Report |
|
27
|
Gemma 4 26B-A4B Instruct (Thinking)
Open Weight
Workstation
vLLM
|
74.9
|
Thinking
|
74Tool-Use supported, Score: 74
|
27.3 t/s
|
–
|
Report |
|
28
|
Claude Haiku 4.5
Commercial
Frontier
API
|
74.8
|
Standard
|
67Tool-Use supported, Score: 67
|
98.4 t/s
|
$0.43
|
Report |
|
29
|
Mistral 3 Large
Open Weight
Server
API
|
74.8
|
Standard
|
69Tool-Use supported, Score: 69
|
50.7 t/s
|
$0.10
|
Report |
|
30
|
Grok 4.6
Commercial
Frontier
API
|
74.7
|
Thinking
|
73Tool-Use supported, Score: 73
|
13.6 t/s
|
$0.44
|
Report |
|
31
|
Xiaomi MiMo V2.5
Open Weight
Server
Cloud
|
74.5
|
Adaptive
|
76Tool-Use supported, Score: 76
|
64.5 t/s
|
$0.04
|
Report |
|
32
|
Ministral 3 14B (Unsloth)
Open Weight
Desktop
llama.cpp
|
74.5
|
Standard
|
61Tool-Use supported, Score: 61
|
15.8 t/s
|
–
|
Report |
|
33
|
GPT-OSS 120B (Thinking)
Open Weight
Server
vLLM
|
74.4
|
Thinking
|
76Tool-Use supported, Score: 76
|
31.4 t/s
|
–
|
Report |
|
34
|
Qwen 3.5 35B-A3B (Unsloth)
Open Weight
Workstation
llama.cpp
|
74.3
|
Adaptive
|
72Tool-Use supported, Score: 72
|
69.3 t/s
|
–
|
Report |
|
35
|
Mistral Medium 3.5
Open Weight
Server
API
|
74.2
|
Standard
|
71Tool-Use supported, Score: 71
|
106 t/s
|
$0.46
|
Report |
|
36
|
Qwen 3 Coder Next
Open Weight
Server
llama.cpp
|
74.1
|
Standard
|
71Tool-Use supported, Score: 71
|
48.5 t/s
|
–
|
Report |
|
37
|
Qwen 3.6 35B-A3B (Unsloth) (Thinking)
Open Weight
Workstation
vLLM
|
74.1
|
Thinking
|
77Tool-Use supported, Score: 77
|
84.6 t/s
|
–
|
Report |
|
38
|
Grok 4.5
Commercial
Frontier
API
|
73.8
|
Standard
|
77Tool-Use supported, Score: 77
|
25.6 t/s
|
$0.52
|
Report |
|
39
|
DeepSeek V4.1 Flash
Open Weight
Server
Cloud
|
73.5
|
Thinking
|
76Tool-Use supported, Score: 76
|
112 t/s
|
$0.12
|
Report |
|
40
|
Qwen 3.6 27B
Open Weight
Workstation
vLLM
|
73.5
|
Standard
|
76Tool-Use supported, Score: 76
|
19.0 t/s
|
–
|
Report |
|
41
|
Qwen 3.8 27B Uncensored
Open Weight
Workstation
vLLM
|
73.4
|
Standard
|
77Tool-Use supported, Score: 77
|
13.4 t/s
|
–
|
Report |
|
42
|
Gemma 4 26B-A4B Q5_K_M (ARA-Abliterated)
Open Weight
Workstation
llama.cpp
|
73.3
|
Standard
|
76Tool-Use supported, Score: 76
|
53.5 t/s
|
–
|
Report |
|
43
|
GLM-5.1
Open Weight
Server
Cloud
|
73.3
|
Adaptive
|
68Tool-Use supported, Score: 68
|
63.5 t/s
|
$0.52
|
Report |
|
44
|
Mistral Small 4
Open Weight
Server
API
|
73.3
|
Standard
|
20Tool-Use supported, Score: 20
|
117 t/s
|
$0.03
|
Report |
|
45
|
Signal 3.8 27B
Open Weight
Workstation
llama.cpp
|
73.2
|
Standard
|
77Tool-Use supported, Score: 77
|
14.5 t/s
|
–
|
Report |
|
46
|
Gemma 4 31B Instruct
Open Weight
Workstation
Cloud
|
73.2
|
Standard
|
74Tool-Use supported, Score: 74
|
30.6 t/s
|
$0.02
|
Report |
|
47
|
Qwen 3.5 397B A17B
Open Weight
Server
Cloud
|
73.2
|
Adaptive
|
74Tool-Use supported, Score: 74
|
20.4 t/s
|
$0.46
|
Report |
|
48
|
Gemini 2.5 Pro
Commercial
Frontier
API
|
73.2
|
Adaptive
|
71Tool-Use supported, Score: 71
|
29.7 t/s
|
$0.68
|
Report |
|
49
|
MiniMax M2.7
Restricted
Server
Cloud
|
73.0
|
Adaptive
|
65Tool-Use supported, Score: 65
|
48.0 t/s
|
$0.13
|
Report |
|
50
|
Kimi K2.7 Code
Open Weight
Frontier
Cloud
|
73.0
|
Thinking
|
66Tool-Use supported, Score: 66
|
100 t/s
|
$0.58
|
Report |
|
51
|
GPT-5.4
Commercial
Frontier
API
|
72.9
|
Standard
|
58Tool-Use supported, Score: 58
|
63.6 t/s
|
$0.94
|
Report |
|
52
|
GLM-4.7
Restricted
Server
Cloud
|
72.9
|
Adaptive
|
71Tool-Use supported, Score: 71
|
58.8 t/s
|
$0.29
|
Report |
|
53
|
Gemma 4 12B Instruct (Unsloth, Q6_K_XL) — Instruct-Profil
Open Weight
Desktop
llama.cpp
|
72.9
|
Standard
|
75Tool-Use supported, Score: 75
|
13.3 t/s
|
–
|
Report |
|
54
|
DeepSeek V4 Flash
Open Weight
Server
Cloud
|
72.9
|
Adaptive
|
78Tool-Use supported, Score: 78
|
29.9 t/s
|
$0.01
|
Report |
|
55
|
GPT-OSS 120B
Open Weight
Server
vLLM
|
72.8
|
Standard
|
67Tool-Use supported, Score: 67
|
30.1 t/s
|
–
|
Report |
|
56
|
Gemma 4 12B Instruct (Unsloth, Q6_K_XL)
Open Weight
Desktop
llama.cpp
|
72.7
|
Thinking
|
71Tool-Use supported, Score: 71
|
16.6 t/s
|
–
|
Report |
|
57
|
DeepSeek V4 Pro
Open Weight
Frontier
Cloud
|
72.6
|
Adaptive
|
77Tool-Use supported, Score: 77
|
53.3 t/s
|
$0.30
|
Report |
|
58
|
NVIDIA Nemotron 3 Ultra 550B A55B
Open Weight
Server
Cloud
|
72.5
|
Adaptive
|
55Tool-Use supported, Score: 55
|
97.0 t/s
|
$0.37
|
Report |
|
59
|
Gemma 4 26B-A4B Instruct
Open Weight
Workstation
vLLM
|
72.0
|
Standard
|
65Tool-Use supported, Score: 65
|
35.2 t/s
|
–
|
Report |
|
60
|
Gemma 4 E4B
Open Weight
Edge
llama.cpp
|
71.9
|
Adaptive
|
73Tool-Use supported, Score: 73
|
48.2 t/s
|
–
|
Report |
|
61
|
Qwen 3.6 35B-A3B (Uncensored)
Open Weight
Workstation
llama.cpp
|
71.7
|
Adaptive
|
60Tool-Use supported, Score: 60
|
53.9 t/s
|
–
|
Report |
|
62
|
Gemini 3.7 Flash
Commercial
Frontier
Cloud
|
71.6
|
Adaptive
|
80Tool-Use supported, Score: 80
|
94.1 t/s
|
–
|
Report |
|
63
|
Muse Glimmer 30B
Open Weight
Workstation
vLLM
|
71.6
|
Thinking
|
73Tool-Use supported, Score: 73
|
10.6 t/s
|
–
|
Report |
|
64
|
Ministral 3 8B (Unsloth)
Open Weight
Edge
llama.cpp
|
71.3
|
Standard
|
60Tool-Use supported, Score: 60
|
24.9 t/s
|
–
|
Report |
|
65
|
Qwen 3.5 27B
Open Weight
Workstation
vLLM
|
71.0
|
Thinking
|
66Tool-Use supported, Score: 66
|
17.8 t/s
|
–
|
Report |
|
66
|
Qwen 3.6 35B-A3B (Unsloth)
Open Weight
Workstation
vLLM
|
70.5
|
Standard
|
56Tool-Use supported, Score: 56
|
67.5 t/s
|
–
|
Report |
|
67
|
Devstral 2
Open Weight
Server
API
|
70.2
|
Standard
|
55Tool-Use supported, Score: 55
|
56.2 t/s
|
$0.13
|
Report |
|
68
|
GPT-5.4 Mini
Commercial
Frontier
API
|
70.2
|
Standard
|
68Tool-Use supported, Score: 68
|
120 t/s
|
$0.25
|
Report |
|
69
|
Qwen 3.5 4B (Unsloth)
Open Weight
Nano
llama.cpp
|
69.9
|
Adaptive
|
78Tool-Use supported, Score: 78
|
51.3 t/s
|
–
|
Report |
|
70
|
NVIDIA Nemotron 3.5 Lightning 30B (Thinking)
Open Weight
Workstation
vLLM
|
69.6
|
Thinking
|
73Tool-Use supported, Score: 73
|
93.5 t/s
|
–
|
Report |
|
71
|
o4-mini
Commercial
Frontier
API
|
69.5
|
Thinking
|
63Tool-Use supported, Score: 63
|
69.7 t/s
|
$0.40
|
Report |
|
72
|
Laguna S 2.1
Open Weight
Server
vLLM
|
69.1
|
Thinking
|
63Tool-Use supported, Score: 63
|
17.2 t/s
|
$0.02
|
Report |
|
73
|
Hermes 4 70B
Open Weight
Server
Cloud
|
69.0
|
Adaptive
|
76Tool-Use supported, Score: 76
|
76.8 t/s
|
$0.03
|
Report |
|
74
|
Gemma 3 12B IT
Restricted
Desktop
llama.cpp
|
68.9
|
Standard
|
64Tool-Use supported, Score: 64
|
39.1 t/s
|
–
|
Report |
|
75
|
Qwen 3.5 9B (Unsloth)
Open Weight
Edge
llama.cpp
|
68.7
|
Adaptive
|
57Tool-Use supported, Score: 57
|
34.3 t/s
|
–
|
Report |
|
76
|
Hermes 4 405B
Restricted
Server
Cloud
|
68.4
|
Adaptive
|
69Tool-Use supported, Score: 69
|
31.1 t/s
|
$0.17
|
Report |
|
77
|
NVIDIA Nemotron 3 Nano 30B A3B
Open Weight
Workstation
Cloud
|
68.3
|
Adaptive
|
66Tool-Use supported, Score: 66
|
35.4 t/s
|
$0.02
|
Report |
|
78
|
Hermes 4 14B (Abliterated)
Open Weight
Desktop
llama.cpp
|
68.2
|
Adaptive
|
70Tool-Use supported, Score: 70
|
25.4 t/s
|
–
|
Report |
|
79
|
Gemma 4 E2B (Unsloth)
Open Weight
Edge
llama.cpp
|
68.1
|
Thinking
|
59Tool-Use supported, Score: 59
|
67.1 t/s
|
–
|
Report |
|
80
|
GPT-5.4 Nano
Commercial
Frontier
API
|
68.1
|
Standard
|
49Tool-Use supported, Score: 49
|
125 t/s
|
$0.07
|
Report |
|
81
|
o3-mini
Commercial
Frontier
API
|
68.0
|
Thinking
|
73Tool-Use supported, Score: 73
|
71.2 t/s
|
$0.37
|
Report |
|
82
|
Hermes 4.3 36B (Thinking)
Open Weight
Server
vLLM
|
68.0
|
Thinking
|
67Tool-Use supported, Score: 67
|
13.3 t/s
|
–
|
Report |
|
83
|
GLM-5.3
Open Weight
Server
Cloud
|
67.7
|
Thinking
|
80Tool-Use supported, Score: 80
|
53.7 t/s
|
$1.47
|
Report |
|
84
|
Qwen 3 14B
Open Weight
Desktop
llama.cpp
|
67.7
|
Adaptive
|
66Tool-Use supported, Score: 66
|
23.9 t/s
|
–
|
Report |
|
85
|
Hermes 4 14B
Open Weight
Desktop
llama.cpp
|
67.3
|
Adaptive
|
67Tool-Use supported, Score: 67
|
30.3 t/s
|
–
|
Report |
|
86
|
Qwen3.8-2.4T-A95B
Open Weight
Frontier
Cloud
|
67.2
|
Thinking
|
78Tool-Use supported, Score: 78
|
44.2 t/s
|
$1.89
|
Report |
|
87
|
Hermes 4.3 36B
Open Weight
Server
vLLM
|
67.1
|
Standard
|
67Tool-Use supported, Score: 67
|
12.7 t/s
|
–
|
Report |
|
88
|
Ornith 1.0 9B (Unsloth)
Open Weight
Desktop
llama.cpp
|
66.8
|
Thinking
|
71Tool-Use supported, Score: 71
|
25.1 t/s
|
–
|
Report |
|
89
|
Codestral 25.08
Restricted
Desktop
API
|
65.5
|
Standard
|
69Tool-Use supported, Score: 69
|
192 t/s
|
$0.05
|
Report |
|
90
|
Ministral 3 3B (Unsloth)
Open Weight
Nano
llama.cpp
|
64.8
|
Standard
|
56Tool-Use supported, Score: 56
|
53.0 t/s
|
–
|
Report |
|
91
|
GPT-OSS 20B
Open Weight
Desktop
vLLM
|
64.3
|
Standard
|
28Tool-Use supported, Score: 28
|
41.5 t/s
|
–
|
Report |
|
92
|
Llama 8B (Unsloth, provenance unverified)
Restricted
Edge
llama.cpp
|
62.7
|
Standard
|
No Tool-Use
|
25.5 t/s
|
–
|
Report |
|
93
|
GPT-OSS 20B (Thinking)
Open Weight
Desktop
vLLM
|
61.6
|
Thinking
|
26Tool-Use supported, Score: 26
|
41.2 t/s
|
–
|
Report |
|
94
|
Gemma 3 4B (Unsloth)
Restricted
Nano
llama.cpp
|
61.5
|
Standard
|
No Tool-Use
|
44.5 t/s
|
–
|
Report |
|
95
|
Qwen 3 4B
Open Weight
Nano
llama.cpp
|
60.9
|
Adaptive
|
71Tool-Use supported, Score: 71
|
74.0 t/s
|
–
|
Report |
|
96
|
Hermes 3 8B
Restricted
Edge
llama.cpp
|
58.8
|
Standard
|
56Tool-Use supported, Score: 56
|
48.0 t/s
|
–
|
Report |
|
97
|
Phi-4 Mini (Unsloth)
Open Weight
Nano
llama.cpp
|
58.4
|
Standard
|
66Tool-Use supported, Score: 66
|
46.1 t/s
|
–
|
Report |
|
98
|
DeepSeek R1 Distill Qwen 14B
Open Weight
Desktop
llama.cpp
|
58.3
|
Standard
|
No Tool-Use
|
15.0 t/s
|
–
|
Report |
|
99
|
Qwen 2.5 Coder 7B
Open Weight
Edge
llama.cpp
|
55.3
|
Standard
|
58Tool-Use supported, Score: 58
|
37.4 t/s
|
–
|
Report |
|
100
|
Llama 3.2 3B (Unsloth)
Restricted
Nano
llama.cpp
|
52.0
|
Standard
|
48Tool-Use supported, Score: 48
|
55.0 t/s
|
–
|
Report |
|
101
|
Llama 3.2 1B (Unsloth)
Restricted
Nano
llama.cpp
|
44.4
|
Standard
|
No Tool-Use
|
130 t/s
|
–
|
Report |
|
102
|
DeepSeek R1 Distill Qwen 7B
Open Weight
Edge
llama.cpp
|
42.6
|
Standard
|
No Tool-Use
|
29.3 t/s
|
–
|
Report |
|
103
|
DeepSeek R1 Distill Qwen 1.5B
Open Weight
Nano
llama.cpp
|
35.4
|
Standard
|
No Tool-Use
|
110 t/s
|
–
|
Report |
|
104
|
Gemma 3 270M (Unsloth)
Restricted
Nano
llama.cpp
|
23.6
|
Standard
|
No Tool-Use
|
211 t/s
|
–
|
Report |