LLMs head to head 7 use-case scenarios · Equal conditions · one results list

How well do open-weight models hold up against proprietary API models and license-restricted restricted-weight models? Using everyday tasks under identical conditions, all license models and weight classes land in the same scoreboard.

The total score shows the best all-rounder. The profiles below answer the real everyday question: is this model sufficient for your task — even if it's a small, local specialist rather than a cloud giant?

Scoreboard

104 models in the scoreboard

# Report
1
GLM-5.3-Flash
Open Weight Server Cloud
80.6
Thinking
78Tool-Use supported, Score: 78
52.7 t/s
$0.11
Report
2
MiniMax M3
Open Weight Server Cloud
80.2
Adaptive
81Tool-Use supported, Score: 81
108 t/s
$0.16
Report
3
Qwen 3.8 27B (Thinking)
Open Weight Workstation vLLM
80.0
Thinking
87Tool-Use supported, Score: 87
21.6 t/s
Report
4
Claude Sonnet 4.6
Commercial Frontier API
79.8
Adaptive
75Tool-Use supported, Score: 75
51.1 t/s
$1.77
Report
5
Qwen3.8-Flash
Open Weight Server Cloud
79.5
Thinking
76Tool-Use supported, Score: 76
50.9 t/s
$0.05
Report
6
Claude Sonnet 5
Commercial Frontier API
79.0
Thinking
80Tool-Use supported, Score: 80
82.5 t/s
$1.43
Report
7
Xiaomi MiMo V2.5 Pro
Open Weight Frontier Cloud
79.0
Adaptive
74Tool-Use supported, Score: 74
48.8 t/s
$0.12
Report
8
Kimi K3
Open Weight Frontier Cloud
78.8
Thinking
81Tool-Use supported, Score: 81
48.3 t/s
$3.30
Report
9
Claude Opus 5
Commercial Frontier API
78.8
Thinking
80Tool-Use supported, Score: 80
63.8 t/s
$4.62
Report
10
Claude Opus 4.8
Commercial Frontier API
77.9
Thinking
72Tool-Use supported, Score: 72
63.3 t/s
$2.71
Report
11
Swift Qwen 3.8 27B
Open Weight Workstation vLLM
77.8
Thinking
74Tool-Use supported, Score: 74
15.5 t/s
Report
12
Gemini 3.5 Flash
Commercial Frontier API
77.2
Adaptive
76Tool-Use supported, Score: 76
58.7 t/s
$0.49
Report
13
Meta Muse Spark 1.2
Commercial Frontier Cloud
76.7
Adaptive
70Tool-Use supported, Score: 70
115 t/s
$0.54
Report
14
Kimi K2.5
Open Weight Frontier Cloud
76.6
Thinking
77Tool-Use supported, Score: 77
37.7 t/s
$0.39
Report
15
GPT-5.5
Commercial Frontier API
76.5
Thinking
79Tool-Use supported, Score: 79
59.3 t/s
$2.97
Report
16
GLM 4.6
Restricted Server Cloud
76.2
Standard
53Tool-Use supported, Score: 53
46.1 t/s
$0.25
Report
17
Qwen 3.8 27B Uncensored (Thinking)
Open Weight Workstation vLLM
76.1
Thinking
74Tool-Use supported, Score: 74
16.1 t/s
Report
18
Ornith 1.5 35B-A3B
Open Weight Workstation vLLM
76.0
Thinking
76Tool-Use supported, Score: 76
61.0 t/s
Report
19
Qwen 3.7 Max
Commercial Frontier Cloud
75.8
Adaptive
60Tool-Use supported, Score: 60
62.6 t/s
$0.75
Report
20
Ornith 1.0 35B
Open Weight Workstation vLLM
75.8
Thinking
63Tool-Use supported, Score: 63
35.5 t/s
Report
21
Qwen 3.6 27B (Thinking)
Open Weight Workstation vLLM
75.6
Thinking
73Tool-Use supported, Score: 73
25.2 t/s
Report
22
Kimi K2.6
Open Weight Frontier Cloud
75.6
Adaptive
74Tool-Use supported, Score: 74
49.4 t/s
$0.96
Report
23
GLM-5.2
Open Weight Server Cloud
75.5
Adaptive
51Tool-Use supported, Score: 51
65.7 t/s
$0.40
Report
24
Qwen 3.6 Plus
Commercial Frontier Cloud
75.5
Adaptive
68Tool-Use supported, Score: 68
56.3 t/s
$0.41
Report
25
Upstage Solar Pro4
Commercial Frontier Cloud
75.4
Standard
71Tool-Use supported, Score: 71
35.9 t/s
$0.01
Report
26
Qwen 3.8 27B
Open Weight Workstation vLLM
75.2
Standard
70Tool-Use supported, Score: 70
17.4 t/s
Report
27
Gemma 4 26B-A4B Instruct (Thinking)
Open Weight Workstation vLLM
74.9
Thinking
74Tool-Use supported, Score: 74
27.3 t/s
Report
28
Claude Haiku 4.5
Commercial Frontier API
74.8
Standard
67Tool-Use supported, Score: 67
98.4 t/s
$0.43
Report
29
Mistral 3 Large
Open Weight Server API
74.8
Standard
69Tool-Use supported, Score: 69
50.7 t/s
$0.10
Report
30
Grok 4.6
Commercial Frontier API
74.7
Thinking
73Tool-Use supported, Score: 73
13.6 t/s
$0.44
Report
31
Xiaomi MiMo V2.5
Open Weight Server Cloud
74.5
Adaptive
76Tool-Use supported, Score: 76
64.5 t/s
$0.04
Report
32
Ministral 3 14B (Unsloth)
Open Weight Desktop llama.cpp
74.5
Standard
61Tool-Use supported, Score: 61
15.8 t/s
Report
33
GPT-OSS 120B (Thinking)
Open Weight Server vLLM
74.4
Thinking
76Tool-Use supported, Score: 76
31.4 t/s
Report
34
Qwen 3.5 35B-A3B (Unsloth)
Open Weight Workstation llama.cpp
74.3
Adaptive
72Tool-Use supported, Score: 72
69.3 t/s
Report
35
Mistral Medium 3.5
Open Weight Server API
74.2
Standard
71Tool-Use supported, Score: 71
106 t/s
$0.46
Report
36
Qwen 3 Coder Next
Open Weight Server llama.cpp
74.1
Standard
71Tool-Use supported, Score: 71
48.5 t/s
Report
37
Qwen 3.6 35B-A3B (Unsloth) (Thinking)
Open Weight Workstation vLLM
74.1
Thinking
77Tool-Use supported, Score: 77
84.6 t/s
Report
38
Grok 4.5
Commercial Frontier API
73.8
Standard
77Tool-Use supported, Score: 77
25.6 t/s
$0.52
Report
39
DeepSeek V4.1 Flash
Open Weight Server Cloud
73.5
Thinking
76Tool-Use supported, Score: 76
112 t/s
$0.12
Report
40
Qwen 3.6 27B
Open Weight Workstation vLLM
73.5
Standard
76Tool-Use supported, Score: 76
19.0 t/s
Report
41
Qwen 3.8 27B Uncensored
Open Weight Workstation vLLM
73.4
Standard
77Tool-Use supported, Score: 77
13.4 t/s
Report
42
Gemma 4 26B-A4B Q5_K_M (ARA-Abliterated)
Open Weight Workstation llama.cpp
73.3
Standard
76Tool-Use supported, Score: 76
53.5 t/s
Report
43
GLM-5.1
Open Weight Server Cloud
73.3
Adaptive
68Tool-Use supported, Score: 68
63.5 t/s
$0.52
Report
44
Mistral Small 4
Open Weight Server API
73.3
Standard
20Tool-Use supported, Score: 20
117 t/s
$0.03
Report
45
Signal 3.8 27B
Open Weight Workstation llama.cpp
73.2
Standard
77Tool-Use supported, Score: 77
14.5 t/s
Report
46
Gemma 4 31B Instruct
Open Weight Workstation Cloud
73.2
Standard
74Tool-Use supported, Score: 74
30.6 t/s
$0.02
Report
47
Qwen 3.5 397B A17B
Open Weight Server Cloud
73.2
Adaptive
74Tool-Use supported, Score: 74
20.4 t/s
$0.46
Report
48
Gemini 2.5 Pro
Commercial Frontier API
73.2
Adaptive
71Tool-Use supported, Score: 71
29.7 t/s
$0.68
Report
49
MiniMax M2.7
Restricted Server Cloud
73.0
Adaptive
65Tool-Use supported, Score: 65
48.0 t/s
$0.13
Report
50
Kimi K2.7 Code
Open Weight Frontier Cloud
73.0
Thinking
66Tool-Use supported, Score: 66
100 t/s
$0.58
Report
51
GPT-5.4
Commercial Frontier API
72.9
Standard
58Tool-Use supported, Score: 58
63.6 t/s
$0.94
Report
52
GLM-4.7
Restricted Server Cloud
72.9
Adaptive
71Tool-Use supported, Score: 71
58.8 t/s
$0.29
Report
53
Gemma 4 12B Instruct (Unsloth, Q6_K_XL) — Instruct-Profil
Open Weight Desktop llama.cpp
72.9
Standard
75Tool-Use supported, Score: 75
13.3 t/s
Report
54
DeepSeek V4 Flash
Open Weight Server Cloud
72.9
Adaptive
78Tool-Use supported, Score: 78
29.9 t/s
$0.01
Report
55
GPT-OSS 120B
Open Weight Server vLLM
72.8
Standard
67Tool-Use supported, Score: 67
30.1 t/s
Report
56
Gemma 4 12B Instruct (Unsloth, Q6_K_XL)
Open Weight Desktop llama.cpp
72.7
Thinking
71Tool-Use supported, Score: 71
16.6 t/s
Report
57
DeepSeek V4 Pro
Open Weight Frontier Cloud
72.6
Adaptive
77Tool-Use supported, Score: 77
53.3 t/s
$0.30
Report
58
NVIDIA Nemotron 3 Ultra 550B A55B
Open Weight Server Cloud
72.5
Adaptive
55Tool-Use supported, Score: 55
97.0 t/s
$0.37
Report
59
Gemma 4 26B-A4B Instruct
Open Weight Workstation vLLM
72.0
Standard
65Tool-Use supported, Score: 65
35.2 t/s
Report
60
Gemma 4 E4B
Open Weight Edge llama.cpp
71.9
Adaptive
73Tool-Use supported, Score: 73
48.2 t/s
Report
61
Qwen 3.6 35B-A3B (Uncensored)
Open Weight Workstation llama.cpp
71.7
Adaptive
60Tool-Use supported, Score: 60
53.9 t/s
Report
62
Gemini 3.7 Flash
Commercial Frontier Cloud
71.6
Adaptive
80Tool-Use supported, Score: 80
94.1 t/s
Report
63
Muse Glimmer 30B
Open Weight Workstation vLLM
71.6
Thinking
73Tool-Use supported, Score: 73
10.6 t/s
Report
64
Ministral 3 8B (Unsloth)
Open Weight Edge llama.cpp
71.3
Standard
60Tool-Use supported, Score: 60
24.9 t/s
Report
65
Qwen 3.5 27B
Open Weight Workstation vLLM
71.0
Thinking
66Tool-Use supported, Score: 66
17.8 t/s
Report
66
Qwen 3.6 35B-A3B (Unsloth)
Open Weight Workstation vLLM
70.5
Standard
56Tool-Use supported, Score: 56
67.5 t/s
Report
67
Devstral 2
Open Weight Server API
70.2
Standard
55Tool-Use supported, Score: 55
56.2 t/s
$0.13
Report
68
GPT-5.4 Mini
Commercial Frontier API
70.2
Standard
68Tool-Use supported, Score: 68
120 t/s
$0.25
Report
69
Qwen 3.5 4B (Unsloth)
Open Weight Nano llama.cpp
69.9
Adaptive
78Tool-Use supported, Score: 78
51.3 t/s
Report
70
NVIDIA Nemotron 3.5 Lightning 30B (Thinking)
Open Weight Workstation vLLM
69.6
Thinking
73Tool-Use supported, Score: 73
93.5 t/s
Report
71
o4-mini
Commercial Frontier API
69.5
Thinking
63Tool-Use supported, Score: 63
69.7 t/s
$0.40
Report
72
Laguna S 2.1
Open Weight Server vLLM
69.1
Thinking
63Tool-Use supported, Score: 63
17.2 t/s
$0.02
Report
73
Hermes 4 70B
Open Weight Server Cloud
69.0
Adaptive
76Tool-Use supported, Score: 76
76.8 t/s
$0.03
Report
74
Gemma 3 12B IT
Restricted Desktop llama.cpp
68.9
Standard
64Tool-Use supported, Score: 64
39.1 t/s
Report
75
Qwen 3.5 9B (Unsloth)
Open Weight Edge llama.cpp
68.7
Adaptive
57Tool-Use supported, Score: 57
34.3 t/s
Report
76
Hermes 4 405B
Restricted Server Cloud
68.4
Adaptive
69Tool-Use supported, Score: 69
31.1 t/s
$0.17
Report
77
NVIDIA Nemotron 3 Nano 30B A3B
Open Weight Workstation Cloud
68.3
Adaptive
66Tool-Use supported, Score: 66
35.4 t/s
$0.02
Report
78
Hermes 4 14B (Abliterated)
Open Weight Desktop llama.cpp
68.2
Adaptive
70Tool-Use supported, Score: 70
25.4 t/s
Report
79
Gemma 4 E2B (Unsloth)
Open Weight Edge llama.cpp
68.1
Thinking
59Tool-Use supported, Score: 59
67.1 t/s
Report
80
GPT-5.4 Nano
Commercial Frontier API
68.1
Standard
49Tool-Use supported, Score: 49
125 t/s
$0.07
Report
81
o3-mini
Commercial Frontier API
68.0
Thinking
73Tool-Use supported, Score: 73
71.2 t/s
$0.37
Report
82
Hermes 4.3 36B (Thinking)
Open Weight Server vLLM
68.0
Thinking
67Tool-Use supported, Score: 67
13.3 t/s
Report
83
GLM-5.3
Open Weight Server Cloud
67.7
Thinking
80Tool-Use supported, Score: 80
53.7 t/s
$1.47
Report
84
Qwen 3 14B
Open Weight Desktop llama.cpp
67.7
Adaptive
66Tool-Use supported, Score: 66
23.9 t/s
Report
85
Hermes 4 14B
Open Weight Desktop llama.cpp
67.3
Adaptive
67Tool-Use supported, Score: 67
30.3 t/s
Report
86
Qwen3.8-2.4T-A95B
Open Weight Frontier Cloud
67.2
Thinking
78Tool-Use supported, Score: 78
44.2 t/s
$1.89
Report
87
Hermes 4.3 36B
Open Weight Server vLLM
67.1
Standard
67Tool-Use supported, Score: 67
12.7 t/s
Report
88
Ornith 1.0 9B (Unsloth)
Open Weight Desktop llama.cpp
66.8
Thinking
71Tool-Use supported, Score: 71
25.1 t/s
Report
89
Codestral 25.08
Restricted Desktop API
65.5
Standard
69Tool-Use supported, Score: 69
192 t/s
$0.05
Report
90
Ministral 3 3B (Unsloth)
Open Weight Nano llama.cpp
64.8
Standard
56Tool-Use supported, Score: 56
53.0 t/s
Report
91
GPT-OSS 20B
Open Weight Desktop vLLM
64.3
Standard
28Tool-Use supported, Score: 28
41.5 t/s
Report
92
Llama 8B (Unsloth, provenance unverified)
Restricted Edge llama.cpp
62.7
Standard
No Tool-Use
25.5 t/s
Report
93
GPT-OSS 20B (Thinking)
Open Weight Desktop vLLM
61.6
Thinking
26Tool-Use supported, Score: 26
41.2 t/s
Report
94
Gemma 3 4B (Unsloth)
Restricted Nano llama.cpp
61.5
Standard
No Tool-Use
44.5 t/s
Report
95
Qwen 3 4B
Open Weight Nano llama.cpp
60.9
Adaptive
71Tool-Use supported, Score: 71
74.0 t/s
Report
96
Hermes 3 8B
Restricted Edge llama.cpp
58.8
Standard
56Tool-Use supported, Score: 56
48.0 t/s
Report
97
Phi-4 Mini (Unsloth)
Open Weight Nano llama.cpp
58.4
Standard
66Tool-Use supported, Score: 66
46.1 t/s
Report
98
DeepSeek R1 Distill Qwen 14B
Open Weight Desktop llama.cpp
58.3
Standard
No Tool-Use
15.0 t/s
Report
99
Qwen 2.5 Coder 7B
Open Weight Edge llama.cpp
55.3
Standard
58Tool-Use supported, Score: 58
37.4 t/s
Report
100
Llama 3.2 3B (Unsloth)
Restricted Nano llama.cpp
52.0
Standard
48Tool-Use supported, Score: 48
55.0 t/s
Report
101
Llama 3.2 1B (Unsloth)
Restricted Nano llama.cpp
44.4
Standard
No Tool-Use
130 t/s
Report
102
DeepSeek R1 Distill Qwen 7B
Open Weight Edge llama.cpp
42.6
Standard
No Tool-Use
29.3 t/s
Report
103
DeepSeek R1 Distill Qwen 1.5B
Open Weight Nano llama.cpp
35.4
Standard
No Tool-Use
110 t/s
Report
104
Gemma 3 270M (Unsloth)
Restricted Nano llama.cpp
23.6
Standard
No Tool-Use
211 t/s
Report