Llama 3.3 70B Versatile
Llama 3.3 70B Versatile is Meta’s all-round model in the 70-billion class, with balanced strengths across a broad range of tasks. With a 128,000-token context window and Open Weights under the Llama 3.3 Community License, the model is available either locally for maximum data sovereignty or through cloud providers.
- Restricted Weights
- Server
- GR
- Text
- Instruction-Tuned
- Real-Time
Sovereign Risk: MEDIUM Meta is a US company; weights are publicly available, and local deployment avoids API data leakage.
Key metrics
Score · Latency · Cost · Quality
- Total Score Bronze
- 64.32
- Routine
- 39.42
- Reasoning
- 24.9
- LLM Judge Avg
- 3.24 / 5
- 100 Coverage
- Avg Task Duration
- 1.75s
- Real-Time
- Token Rate
- 275.73tok/s
- Output Rate
- P95 Latency
- 3.32s
- Top 5 %
- Total Tokens
- 40800
- Output Volume
- Cost per 1K
- $0.0008
- USD / 1K Requests
- Benchmark Cost
- $0.03
- Total · 40800 tok
Benchmark modules
10 modules · weighted · vs. model median & top performer
Token efficiency & latency
Consumption per module vs. model median