Llama 3.3 Nemotron Super 49B v1.5
NVIDIA Llama 3.3 Nemotron Super 49B v1.5 is a pruning- and distillation-optimized variant of Meta’s Llama 3.3 70B with 49 billion parameters. The model delivers strong reasoning performance at reduced resource requirements, a context window of 131,000 tokens, and an optional thinking mode controlled via system prompt. Available as an Open Weights variant under the NVIDIA Open Model License, locally or through cloud providers.
- Open Weights
- Server
- OR
- Text
- Instruction-Tuned
- Interactive
Key metrics
Score · Latency · Cost · Quality
- Total Score Silver
- 70.2
- Routine
- 42.89
- Reasoning
- 27.31
- LLM Judge Avg
- 3.49 / 5
- 100 Coverage
- Avg Task Duration
- 36.91s
- Interactive
- Token Rate
- 20.73tok/s
- Output Rate
- P95 Latency
- 83.14s
- Top 5 %
- Total Tokens
- 105100
- Output Volume
- Cost per 1K
- $0.0004
- USD / 1K Requests
- Benchmark Cost
- $0.04
- Total · 105100 tok
Benchmark modules
10 modules · weighted · vs. model median & top performer
Token efficiency & latency
Consumption per module vs. model median