GPT-OSS 20B

What most OpenAI models can’t do: GPT-OSS 20B is OpenAI’s first Open Weights release since GPT-2 (August 5, 2025) under the Apache 2.0 license. The MoE with 21 billion total and 3.6 billion active parameters runs on a single consumer GPU with only around 16 GB of memory thanks to native MXFP4 quantization, supports tool use via the Harmony format, and offers 131,072 tokens of context as well as configurable reasoning intensity (low/medium/high).

OpenAI Version 1.0 Commercial use permitted MoE 21 B (3.6 B active) 128 K Context 06/2024 locally tested

  • Open Weights
  • Desktop
  • vLLM
  • Text
  • Long Context
  • Interactive

Sovereign Risk: LOW OpenAI is a US company; the model is released as Open Weights under Apache-2.0. Local deployment completely eliminates any API data leakage to OpenAI servers, which is why the risk is rated as low despite US jurisdiction (CLOUD Act).

Key metrics

Score · Latency · Cost · Quality

Total Score Bronze
64.31
Routine
39.51
Reasoning
24.8

Rank #85

LLM Judge Avg
3.6
100 Coverage
Avg Task Duration
26.97
Interactive
Token Rate
41.46
Output Rate
P95 Latency
74.51
Top 5 %
Total Tokens
85300
Output Volume
Cost per 1K
$0
USD / 1K Requests
Benchmark Cost
$0
Total · 85300 tok

Benchmark modules

10 modules · weighted · vs. model median & top performer

GPT-OSS 20B Best model Ø All models
Code Quality 78.9
CLI Benchmark 87.22
Logical Reasoning 66.24
UX Writing 69.15
Documentation 61.16
Content Transform. 66.55
Cultural Intelligence 69
Synthesis Quality 20
Tool Execution 35
ToolUse Score 27.75
Benchmark Cost $0

Token efficiency & latency

Consumption per module vs. model median

Token consumption per module

Performance profile