Hermes 4 405B

Hermes 4 405B is a high-performance instruct and reasoning model from Nous Research with 405 billion parameters, designed for complex reasoning tasks and agentic workflows. The model supports optional thinking, precise tool calls, and structured outputs. Trained for high steerability and reduced Refusal rates. Available as an Open Weights model under the Meta Llama Community License.

NousResearch Version 4 Commercial use permitted Dense 405 B (405 B active) 128 K Context 01/2025 $1 / $3 per 1M

  • Restricted Weights
  • Frontier
  • OR
  • Text
  • Instruction-Tuned
  • Real-Time

Sovereign Risk: LOW Nous Research is a US-based company and subject to the CLOUD Act; however, the weights are publicly available and can be run locally, so no third-party API access is required.

Key metrics

Score · Latency · Cost · Quality

Total Score Silver
67.76
Routine
40.87
Reasoning
26.9

Rank #81

LLM Judge Avg
3.3
100 Coverage
Avg Task Duration
18.37
Real-Time
Token Rate
39.49
Output Rate
P95 Latency
37.76
Top 5 %
Total Tokens
48900
Output Volume
Cost per 1K
$0.003
USD / 1K Requests
Benchmark Cost
$0.15
Total · 48900 tok

Benchmark modules

10 modules · weighted · vs. model median & top performer

Hermes 4 405B Best model Ø All models
Code Quality 67.64
CLI Benchmark 75.67
Logical Reasoning 66.26
UX Writing 68.93
Documentation 63.73
Content Transform. 70.49
Cultural Intelligence 61.04
Synthesis Quality 60
Tool Execution 90
ToolUse Score 74.33
Benchmark Cost $0.15

Token efficiency & latency

Consumption per module vs. model median

Token consumption per module

Performance profile