Hermes 3 8B
Hermes 3 8B is an uncensored fine-tune by NousResearch based on Meta’s Llama 3.1 8B. With eight billion parameters and a 128,000-token context window, the model targets instruction following, tool use, and creative or ambiguous requests. Deployable locally under the Llama 3.1 Community License, the reduced refusal rate is a defining characteristic of this distribution.
- Restricted Weights
- Edge
- llama.cpp
- Text
- Instruction-Tuned
- Uncensored
- Real-Time
Sovereign Risk: MEDIUM NousResearch is a US-based company; the CLOUD Act is only relevant when using the API, not when running the Open Weights variant locally.
Key metrics
Score · Latency · Cost · Quality
- Total Score Bronze
- 58.83
- Routine
- 36.48
- Reasoning
- 22.35
- LLM Judge Avg
- 2.78 / 5
- 100 Coverage
- Avg Task Duration
- 12.93s
- Real-Time
- Token Rate
- 47.99tok/s
- Output Rate
- P95 Latency
- 28.73s
- Top 5 %
- Total Tokens
- 44400
- Output Volume
- Cost per 1K
- $0
- USD / 1K Requests
- Benchmark Cost
- $0
- Total · 44400 tok
Benchmark modules
10 modules · weighted · vs. model median & top performer
Token efficiency & latency
Consumption per module vs. model median