Hermes 3 8B

Hermes 3 8B is an uncensored fine-tune by NousResearch based on Meta’s Llama 3.1 8B. With eight billion parameters and a 128,000-token context window, the model targets instruction following, tool use, and creative or ambiguous requests. Deployable locally under the Llama 3.1 Community License, the reduced refusal rate is a defining characteristic of this distribution.

NousResearch Version 3 Commercial use permitted Dense 8 B (8 B active) 128 K Context 09/2024 locally tested

  • Restricted Weights
  • Edge
  • llama.cpp
  • Text
  • Instruction-Tuned
  • Uncensored
  • Real-Time

Sovereign Risk: MEDIUM NousResearch is a US-based company; the CLOUD Act is only relevant when using the API, not when running the Open Weights variant locally.

Key metrics

Score · Latency · Cost · Quality

Total Score Bronze
58.83
Routine
36.48
Reasoning
22.35

Rank #89

LLM Judge Avg
2.78
100 Coverage
Avg Task Duration
12.93
Real-Time
Token Rate
47.99
Output Rate
P95 Latency
28.73
Top 5 %
Total Tokens
44400
Output Volume
Cost per 1K
$0
USD / 1K Requests
Benchmark Cost
$0
Total · 44400 tok

Benchmark modules

10 modules · weighted · vs. model median & top performer

Hermes 3 8B Best model Ø All models
Code Quality 56.7
CLI Benchmark 80.56
Logical Reasoning 48.82
UX Writing 59.15
Documentation 48.06
Content Transform. 64.24
Cultural Intelligence 67.6
Synthesis Quality 30
Tool Execution 83.33
ToolUse Score 56.38
Benchmark Cost $0

Token efficiency & latency

Consumption per module vs. model median

Token consumption per module

Performance profile