GPT-OSS 120B (Thinking)

GPT-OSS 120B is OpenAI’s largest Open Weights model as of August 2025, released under the Apache 2.0 license with free commercial use. The MoE architecture bundles 116.8 billion total parameters while activating only 5.1 billion per token, and runs on a single high-memory GPU thanks to native MXFP4 quantization. Three reasoning levels and native tool use in the Harmony format round out the profile.

OpenAI Version 1.0 Commercial use permitted MoE 116.8 B (5.1 B active) 131 K Context locally tested

  • Open Weights
  • Server
  • VSPK
  • Text
  • Configurable-Reasoning
  • MXFP4
  • Native-Quant
  • Harmony
  • Interactive

Sovereign Risk: LOW OpenAI is a US company; the model is released as Open Weights under Apache-2.0. Local deployment completely eliminates any API data transfer to OpenAI servers, which is why the risk is rated as low despite US jurisdiction (CLOUD Act).

Key metrics

Score · Latency · Cost · Quality

Total Score Silver
73.41
Routine
44.77
Reasoning
28.64

Rank #38

LLM Judge Avg
3.65
100 Coverage
Avg Task Duration
36.63
Interactive
Token Rate
26.87
Output Rate
P95 Latency
87.49
Top 5 %
Total Tokens
85300
Output Volume
Cost per 1K
$0
USD / 1K Requests
Benchmark Cost
$0
Total · 85300 tok

Benchmark modules

10 modules · weighted · vs. model median & top performer

GPT-OSS 120B (Thinking) Best model Ø All models
Code Quality 78.6
CLI Benchmark 88.33
Logical Reasoning 68.09
UX Writing 62.25
Documentation 69.68
Content Transform. 76.13
Cultural Intelligence 80
Synthesis Quality 55
Tool Execution 90
ToolUse Score 71.67
Benchmark Cost $0

Token efficiency & latency

Consumption per module vs. model median

Token consumption per module

Performance profile