GPT-OSS 120B
GPT-OSS 120B is OpenAI’s largest Open Weights model as of August 2025, released under the Apache 2.0 license with free commercial use. The MoE architecture bundles 116.8 billion total parameters while activating only 5.1 billion per token, and runs on a single high-memory GPU thanks to native MXFP4 quantization. Three reasoning levels and native tool use in the Harmony format round out the profile.
- Open Weights
- Server
- VSPK
- Text
- Configurable-Reasoning
- MXFP4
- Native-Quant
- Harmony
- Interactive
Sovereign Risk: LOW OpenAI is a US company; the model is released as Open Weights under Apache-2.0. Local deployment completely eliminates any API data transfer to OpenAI servers, which is why the risk is rated as low despite US jurisdiction (CLOUD Act).
Key metrics
Score · Latency · Cost · Quality
- Total Score Silver
- 73.83
- Routine
- 45.59
- Reasoning
- 28.23
- LLM Judge Avg
- 3.65 / 5
- 100 Coverage
- Avg Task Duration
- 36.7s
- Interactive
- Token Rate
- 27.53tok/s
- Output Rate
- P95 Latency
- 100.97s
- Top 5 %
- Total Tokens
- 86000
- Output Volume
- Cost per 1K
- $0
- USD / 1K Requests
- Benchmark Cost
- $0
- Total · 86000 tok
Benchmark modules
10 modules · weighted · vs. model median & top performer
Token efficiency & latency
Consumption per module vs. model median