Hermes 4 405B

Hermes 4 405B is a high-performance instruct and reasoning model from Nous Research with 405 billion parameters, designed for complex reasoning tasks and agentic workflows. The model supports optional thinking, precise tool calls, and structured outputs. Trained for high steerability and reduced Refusal rates. Available as a Restricted Weights model under the Meta Llama Community License.

NousResearch Version 4 Commercial use permitted Dense 405 B 128 K Context 01/2025 $1 / $3 per 1M

  • Restricted Weights
  • Server
  • OpenRouter
  • Text
  • Instruction-Tuned
  • Interactive

Sovereign Risk: LOW Nous Research is a US-based company and subject to the CLOUD Act; however, the weights are publicly available and can be run locally, so no third-party API access is required.

Tool-use profile: 6 assets in detail

Comparison of asset performance (P1/P2/Combined) against the fleet average

Asset performance (radar)

Score Breakdown vs. Fleet Average


Tool-use details

Asset performance, reliability, and runtime profile

CrucibleMark evaluates tool use across 6 independent tests. Click a test name for details.

Reliability

  • Tool Call Valid: No
  • Retry: Not required
  • Hallucination: Detected

Reliability measures how consistently a model actually executes tool calls: Tool Call Valid schema and format accepted, Retry Required successful only after retry, Hallucination Flag fabricated tools or parameters detected. All three green means production-ready.

Runtime profile

Call 1
3.14
First Request
MCP
1.13
Protocol Latency
Synthesis
6.21
Response Generation
Total
62.8
Sum of All Phases
Token
6872
Input + Output
Cost
$0
Cost per Run

The runtime profile shows the latency and cost metrics for the model run: Call 1 First Request, MCP Protocol Latency, Synthesis Response Generation, Total sum of all phases, supplemented by Token input and output and Cost cost per run.

Tool-use review

Updated · Instruction-Tuned

Deployment Verdict

Conditional deploy: tool execution is strong, but one invalid tool call and a detected hallucination limit confidence in production MCP pipelines. The combined score of 68.50 does not support uncontrolled pass-through to critical tool infrastructure.

Tool Execution Profile

Hermes 4 405B has a fundamentally sound grasp of tool selection. In the Web Search & Tool Selection test — which checks whether the model chooses search over direct fetch without an explicit hint — it makes the right call reliably. This argues against a rigid pattern and in favor of genuine situational assessment. In the EU License Research test it also reaches cleanly for external sources rather than answering from training data alone.

Operational precision is weaker. In the URL Construction test, which measures correct derivation of a target URL followed by a fetch, performance is serviceable but not deterministic enough for fragile pipelines. The signal tool_call_valid=false fits this picture: the model is not consistently MCP-protocol-clean. This is not a planning problem — it is an execution risk at the tool boundary. On the positive side, no retry was required, so the behavior appears not unstable but punctually imprecise.

Synthesis Fidelity

How well does it condense tool results? Only moderately. The P2 score of 60 reveals a clear gap between retrieval and processing. Particularly in HTTP Fetch & Extract and Multilingual Search & Synthesis — where exact extraction and cross-lingual condensation are required — the model loses precision. For pure retrieval pipelines this is tolerable. For reports, compliance summaries, or decision-relevant synthesis it is too imprecise.

Does it stay within the tool result or fall back on training data? In the honeypot EU License Research test it formally stays on the correct path and does not hallucinate off the cuff. That is the critical trust signal. At the same time, hallucination_flag=true is a global safety risk: once a model can output fabricated facts as a tool result even in isolated cases, the entire tool chain becomes subject to verification.

Error Resilience

When tools fail, Hermes 4 405B responds in a production-appropriate manner. In the 404 test — which checks for transparent error communication versus hallucinated substitute content — it does not invent page content and communicates the failure cleanly. This is operationally acceptable and considerably more important than elegant phrasing.

Operational Profile

Total 62.80s per run. Call 1: 3.14s, MCP latency: 1.13s, Call 2: 6.21s. Slow for the performance delivered. Cost/run: local, therefore financially inexpensive but infrastructurally heavyweight.

Conclusion & Recommendation

Suitable for agentic retrieval pipelines with clear guardrails, logging, and downstream validation of tool results. Not suitable for compliance, policy, or executive summary pipelines where the first synthesis must already be reliable. If you deploy Hermes 4 405B, treat it as a capable tool user with a controlled output layer — not as a trusted final authority for condensed facts.

This analysis was generated automatically based on benchmark data. Model used: GPT-5.4 by OpenAI. The raw data and full methodology are documented in the GitHub project.