GPT-OSS 20B (Thinking)

What most OpenAI models can’t do: GPT-OSS 20B is OpenAI’s first Open Weights release since GPT-2 (August 5, 2025) under the Apache 2.0 license. The MoE with 21 billion total and 3.6 billion active parameters runs on a single consumer GPU with only around 16 GB of memory thanks to native MXFP4 quantization, supports tool use via the Harmony format, and offers 131,072 tokens of context as well as configurable reasoning intensity (low/medium/high).

OpenAI Version 1.0 Commercial use permitted MoE 21 B (3.6 B active) 128 K Context 06/2024 locally tested

  • Open Weights
  • Desktop
  • vLLM
  • Text
  • Long Context
  • Interactive

Sovereign Risk: LOW OpenAI is a US company; the model is released as Open Weights under Apache-2.0. Local deployment completely eliminates any API data leakage to OpenAI servers, which is why the risk is rated as low despite US jurisdiction (CLOUD Act).

Key metrics

Score · Latency · Cost · Quality

Total Score Bronze
61.59
Routine
37.6
Reasoning
23.99

Rank #87

LLM Judge Avg
3.3
100 Coverage
Avg Task Duration
30.14
Interactive
Token Rate
41.2
Output Rate
P95 Latency
57.26
Top 5 %
Total Tokens
92100
Output Volume
Cost per 1K
$0
USD / 1K Requests
Benchmark Cost
$0
Total · 92100 tok

Benchmark modules

10 modules · weighted · vs. model median & top performer

GPT-OSS 20B (Thinking) Best model Ø All models
Code Quality 66.3
CLI Benchmark 83.89
Logical Reasoning 65.67
UX Writing 65.85
Documentation 60.75
Content Transform. 61.75
Cultural Intelligence 73.6
Synthesis Quality 16.67
Tool Execution 35
ToolUse Score 26.08
Benchmark Cost $0

Token efficiency & latency

Consumption per module vs. model median

Token consumption per module

Performance profile