Hermes 4 405B

Hermes 4 405B is a high-performance instruct and reasoning model from Nous Research with 405 billion parameters, designed for complex reasoning tasks and agentic workflows. The model supports optional thinking, precise tool calls, and structured outputs. Trained for high steerability and reduced Refusal rates. Available as an Open Weights model under the Meta Llama Community License.

NousResearch Version 4 Commercial use permitted Dense 405 B (405 B active) 128 K Context 01/2025 $1 / $3 per 1M

  • Restricted Weights
  • Frontier
  • OR
  • Text
  • Instruction-Tuned
  • Real-Time

Sovereign Risk: LOW Nous Research is a US-based company and subject to the CLOUD Act; however, the weights are publicly available and can be run locally, so no third-party API access is required.

Tool-use profile: 6 assets in detail

Comparison of asset performance (P1/P2/Combined) against the fleet average

Asset performance (radar)

Score Breakdown vs. Fleet Average


Tool-use details

Asset performance, reliability, and runtime profile

CrucibleMark evaluates tool use across 6 independent tests. Click a test name for details.

Reliability

  • Tool Call Valid: No
  • Retry: Not required
  • Hallucination: Not detected

Reliability measures how consistently a model actually executes tool calls: Tool Call Valid schema and format accepted, Retry Required successful only after retry, Hallucination Flag fabricated tools or parameters detected. All three green means production-ready.

Runtime profile

Call 1
1.21
First Request
MCP
0.94
Protocol Latency
Synthesis
4.22
Response Generation
Total
38.22
Sum of All Phases
Token
4458
Input + Output
Cost
$0.0068
Cost per Run

The runtime profile shows the latency and cost metrics for the model run: Call 1 First Request, MCP Protocol Latency, Synthesis Response Generation, Total sum of all phases, supplemented by Token input and output and Cost cost per run.

Tool-use review

· Instruction-Tuned

Deployment Verdict

Created on: 14.06.2026, 16:11:20

Conditional deploy, because Hermes 4 405B serves the tool infrastructure reliably and showed no hallucination during the run, but the condensation of tool results remains too imprecise for production-critical synthesis.

Tool Execution Profile

The model is strong on the execution side. Tool calls were valid, MCP-compliant, and runnable without retry. For a pipeline, that matters more than linguistic elegance. In the Web Search & Tool Selection test — which checks whether web_search is chosen over fetch without an explicit hint — Hermes 4 405B reliably identifies the correct tool class. This argues against a rigid retrieval pattern and in favor of genuine tool selection. In the URL Construction test, which measures the correct target URL derived from model knowledge followed by a subsequent fetch, it remains usable but not deterministic enough for fragile endpoints. The pattern is clear: good selection of tool type, somewhat weaker precision on self-derived target addresses.

Synthesis Fidelity

How well does it condense tool results? Only adequately. A P2 of 60 reveals a model that often carries retrieved information forward correctly, but summarizes too coarsely in several tasks. This is most visible in EU License Research and Multilingual Search & Synthesis, where the retrieval works but the condensation does not reliably preserve important caveats and nuances. For search-then-answer use cases, that is often sufficient. For compliance, policy, or precise decision documents, it is not.

Does it stay within the tool result or fall back on training? Predominantly yes — and that is the more important trust finding. In the honeypot EU License Research test, which checks whether current license restrictions are answered from web sources rather than training knowledge, no hallucination occurred. Despite weak synthesis, the model therefore remained within the retrieved evidence space. That is a solid signal for tool trustworthiness.

Error Resilience

Acceptable for production. In the 404 test — which measures transparent behavior on a failing tool call rather than fabricated page content — Hermes 4 405B communicates the error cleanly and does not hallucinate substitute content. Exactly this behavior protects downstream systems from silent factual errors.

Operational Profile

Total 38.22s per run. Tool call 1.21s, MCP latency 0.94s, second model call 4.22s. Operationally on the slower side. Cost per run 0.006770. For a 405B model, that is inexpensive to very well justified.

Conclusion & Recommendation

Suitable for MCP pipelines where clean tool usage, robust error handling, and open weights matter more than precise result condensation. A good fit for research orchestration, retrieval with human review, and agentic pre-stages. Not the first choice for compliance outputs, multilingual executive summaries, or any pipeline where the answer is used directly as a binding final version. In those cases, a strict verifier or a second synthesis model should be added downstream.

This analysis was generated automatically based on benchmark data. Model used: GPT 4.5 by OpenAI. The raw data and full methodology are documented in the GitHub project.