GPT-OSS 20B (Thinking)
What most OpenAI models can’t do: GPT-OSS 20B is OpenAI’s first Open Weights release since GPT-2 (August 5, 2025) under the Apache 2.0 license. The MoE with 21 billion total and 3.6 billion active parameters runs on a single consumer GPU with only around 16 GB of memory thanks to native MXFP4 quantization, supports tool use via the Harmony format, and offers 131,072 tokens of context as well as configurable reasoning intensity (low/medium/high).
- Open Weights
- Desktop
- vLLM
- Text
- Long Context
- Interactive
Sovereign Risk: LOW OpenAI is a US company; the model is released as Open Weights under Apache-2.0. Local deployment completely eliminates any API data leakage to OpenAI servers, which is why the risk is rated as low despite US jurisdiction (CLOUD Act).
Key metrics
Score · Latency · Cost · Quality
- Total Score Bronze
- 61.59
- Routine
- 37.6
- Reasoning
- 23.99
- LLM Judge Avg
- 3.3 / 5
- 100 Coverage
- Avg Task Duration
- 30.14s
- Interactive
- Token Rate
- 41.2tok/s
- Output Rate
- P95 Latency
- 57.26s
- Top 5 %
- Total Tokens
- 92100
- Output Volume
- Cost per 1K
- $0
- USD / 1K Requests
- Benchmark Cost
- $0
- Total · 92100 tok
Benchmark modules
10 modules · weighted · vs. model median & top performer
Token efficiency & latency
Consumption per module vs. model median