Llama 3.3 Nemotron Super 49B v1.5

NVIDIA Llama 3.3 Nemotron Super 49B v1.5 is a pruning- and distillation-optimized variant of Meta’s Llama 3.3 70B with 49 billion parameters. The model delivers strong reasoning performance at reduced resource requirements, a context window of 131,000 tokens, and an optional thinking mode controlled via system prompt. Available as an Open Weights variant under the NVIDIA Open Model License, locally or through cloud providers.

NVIDIA Version 3.3 Super v1.5 Commercial use permitted Dense 49 B (49 B active) 131 K Context 12/2024 $0.4 / $0.4 per 1M

  • Open Weights
  • Server
  • OR
  • Text
  • Instruction-Tuned
  • Interactive

Key metrics

Score · Latency · Cost · Quality

Total Score Silver
70.2
Routine
42.89
Reasoning
27.31

Rank #64

LLM Judge Avg
3.49
100 Coverage
Avg Task Duration
36.91
Interactive
Token Rate
20.73
Output Rate
P95 Latency
83.14
Top 5 %
Total Tokens
105100
Output Volume
Cost per 1K
$0.0004
USD / 1K Requests
Benchmark Cost
$0.04
Total · 105100 tok

Benchmark modules

10 modules · weighted · vs. model median & top performer

Llama 3.3 Nemotron Super 49B v1.5 Best model Ø All models
Code Quality 66
CLI Benchmark 80.34
Logical Reasoning 66.22
UX Writing 68.69
Documentation 64.72
Content Transform. 76.5
Cultural Intelligence 71.72
Synthesis Quality 55.83
Tool Execution 90
ToolUse Score 72.92
Benchmark Cost $0.04

Token efficiency & latency

Consumption per module vs. model median

Token consumption per module

Performance profile