Llama 4 Scout 17B
Llama 4 Scout is Meta’s multimodal fourth-generation Llama model, combining general language processing with image understanding in an efficient MoE architecture. Of 109 billion total parameters, only 17 billion are active per token; the context window spans 128,000 tokens. Available under the Llama 4 Community License, which contains restrictions for EU-based users regarding self-hosting and deployment.
- Restricted Weights
- Server
- Groq
- Text
- Vision
- Instruction-Tuned
- Real-Time
Sovereign Risk: MEDIUM Meta is a US company and subject to the CLOUD Act, which may allow government access to data when using the API. Weights are publicly available. The Llama 4 Community License excludes multimodal Llama 4 models for EU-domiciled entities with respect to self-hosting/deployment; end-user access via third-party APIs is to be assessed separately.
Key metrics
Score · Latency · Cost · Quality
- Total Score Bronze
- 55.17
- Routine
- 34.16
- Reasoning
- 21.01
- LLM Judge Avg
- 3.17 / 5
- 100 Coverage
- Avg Task Duration
- 1.45s
- Real-Time
- Token Rate
- 354.96tok/s
- Output Rate
- P95 Latency
- 2.88s
- Top 5 %
- Total Tokens
- 38900
- Output Volume
- Cost per 1K
- $0.0003
- USD / 1K Requests
- Benchmark Cost
- $0.01
- Total · 38900 tok
Benchmark modules
10 modules · weighted · vs. model median & top performer
Token efficiency & latency
Consumption per module vs. model median