Tool-use review
· Instruction-Tuned
Deployment Verdict
Created on: 14.06.2026, 16:11:20
Conditional deploy, because Hermes 4 405B serves the tool infrastructure reliably and showed no hallucination during the run, but the condensation of tool results remains too imprecise for production-critical synthesis.
Tool Execution Profile
The model is strong on the execution side. Tool calls were valid, MCP-compliant, and runnable without retry. For a pipeline, that matters more than linguistic elegance. In the Web Search & Tool Selection test — which checks whether web_search is chosen over fetch without an explicit hint — Hermes 4 405B reliably identifies the correct tool class. This argues against a rigid retrieval pattern and in favor of genuine tool selection. In the URL Construction test, which measures the correct target URL derived from model knowledge followed by a subsequent fetch, it remains usable but not deterministic enough for fragile endpoints. The pattern is clear: good selection of tool type, somewhat weaker precision on self-derived target addresses.
Synthesis Fidelity
How well does it condense tool results? Only adequately. A P2 of 60 reveals a model that often carries retrieved information forward correctly, but summarizes too coarsely in several tasks. This is most visible in EU License Research and Multilingual Search & Synthesis, where the retrieval works but the condensation does not reliably preserve important caveats and nuances. For search-then-answer use cases, that is often sufficient. For compliance, policy, or precise decision documents, it is not.
Does it stay within the tool result or fall back on training? Predominantly yes — and that is the more important trust finding. In the honeypot EU License Research test, which checks whether current license restrictions are answered from web sources rather than training knowledge, no hallucination occurred. Despite weak synthesis, the model therefore remained within the retrieved evidence space. That is a solid signal for tool trustworthiness.
Error Resilience
Acceptable for production. In the 404 test — which measures transparent behavior on a failing tool call rather than fabricated page content — Hermes 4 405B communicates the error cleanly and does not hallucinate substitute content. Exactly this behavior protects downstream systems from silent factual errors.
Operational Profile
Total 38.22s per run. Tool call 1.21s, MCP latency 0.94s, second model call 4.22s. Operationally on the slower side. Cost per run 0.006770. For a 405B model, that is inexpensive to very well justified.
Conclusion & Recommendation
Suitable for MCP pipelines where clean tool usage, robust error handling, and open weights matter more than precise result condensation. A good fit for research orchestration, retrieval with human review, and agentic pre-stages. Not the first choice for compliance outputs, multilingual executive summaries, or any pipeline where the answer is used directly as a binding final version. In those cases, a strict verifier or a second synthesis model should be added downstream.
This analysis was generated automatically based on benchmark data. Model used: GPT 4.5 by OpenAI. The raw data and full methodology are documented in the GitHub project.