Hermes 4 70B

Hermes 4 70B is an open instruct and reasoning model by Nous Research from the Hermes 4 family with 70 billion parameters. The model combines optional thinking with advanced tool use and structured outputs, trained for high steerability and reduced Refusal rates. Available as an Open Weights model under a Modified MIT license for local or server-side deployment.

NousResearch Version 4 Commercial use permitted Dense 70 B (70 B active) 131 K Context 01/2025 $0.13 / $0.4 per 1M

  • Open Weights
  • Server
  • OR
  • Text
  • Instruction-Tuned
  • Real-Time

LLM Model Review

· Instruction-Tuned

With an overall score of 70.68 percent, Hermes 4 70B presents a clear profile: fast, steerable, often surprisingly capable, but lacking the authority one would reasonably expect from a Server-class model with 70 billion dense parameters and an explicit reasoning focus. The Speed Profile badge “Real-Time Tool Expert” fits well: this model responds quickly, structures output cleanly, and stays close to production-ready — just not always with the last degree of precision. As a Cloud Open Weights model, Hermes 4 70B ran here via OpenRouter; the measured 84.61 tokens per second therefore reflects the performance of OpenRouter’s cloud infrastructure first and foremost, not some universal law of the model itself. Sovereign Risk: MEDIUM — the weights originate from Nous Research in the US; for European users running the model via US-adjacent infrastructure, questions of jurisdiction and potential access under the CLOUD Act remain real.

Header Grades: Stability and Reliability

Metric Value Rating Analysis
Timeout Rate 0/43 Stable The model ran completely stable and reliable throughout testing.
P95 Response Time 24.88 s Consistent Very low tail, almost no outliers.

That’s the good news first. Hermes 4 70B produced zero timeouts across the entire run. For a cloud model of this size class, that’s not a minor detail — it’s the entry ticket to any serious pipeline. Anyone working with agents, batch jobs, or automated review steps doesn’t need genius with gaps; they need a model that simply shows up when called. That’s exactly what Hermes does here.

The latency side also looks pleasingly disciplined. In five percent of all requests, response time was 24.88 seconds or above. That’s not spectacularly fast in the sense of a chat experience without noticeable wait times, but it’s clearly in the green zone. This finding is particularly relevant because Hermes 4 70B belongs to the metadata category Thinking-Optional: the benchmark was deliberately run without Extended Thinking activated. That means what was measured is the default behavior any regular API user gets out of the box. And that default behavior is fast enough to stay out of the way.

Architecture and Character: Instruct First, Reasoning on Demand

The metadata combination Instruct, Thinking-Optional describes Hermes 4 70B with surprising accuracy. The model clearly wants to comply first and impress second. It answers concisely, keeps formats clean most of the time, rarely wastes tokens, and doesn’t exhibit the self-indulgent over-explanation that some reasoning models deploy on every task. At the same time, it’s officially classified for reasoning — not as a mere chat generalist, but as a model from which one should expect more than solid average performance in logic, analysis, and multi-step tasks.

That’s precisely where the test’s tension originates. As a Dense model, all 70 billion parameters are active on every response. In other words, there’s no mixture-of-experts trick here simulating efficiency through a handful of active sub-networks. Hermes 4 70B competes at full capacity. That makes strong average performance respectable — but it also means weaknesses in core areas can’t simply be explained away. A Server-class model of this caliber shouldn’t merely be “quite decent” in code, logic, and documentation. It should project authority. Hermes 4 70B does that sometimes. It doesn’t do it consistently.

Performance Profile: Fast, Cheap, Remarkably Token-Efficient

At 84.61 tokens per second, Hermes 4 70B is unusually fast for a large Open Weights model in the cloud. Combined with a benchmark cost of $0.0181 at $0.0004 per 1K tokens, the result is an economically attractive profile. The badge “Real-Time Tool Expert” signals exactly that: this model is designed for operational, direct use rather than contemplative long-haul tasks. It’s meant to respond, structure, support tools, and not become a bottleneck in real time.

The infrastructure context matters here. Hermes 4 70B did not run as a proprietary full-service offering but as a Cloud Open Weights model via OpenRouter. Speed is therefore always partly a benchmark result of the specific provider, its routing, and its compute environment. For readers, this means: the 84.61 tokens per second is a very solid real-world value for this endpoint. It is not a universal fingerprint to be transferred blindly to any other deployment.

Token economy is another positive. No module exceeds the expected verbosity range. In fact, Reasoning averages 552 output tokens — well below the fleet median of 1,174 — CLI comes in at 176 versus 287, Code Quality at 1,581 versus 2,317, and Documentation at 2,189 versus 2,838. Hermes 4 70B behaves token-efficiently. It rarely talks longer than necessary. On a cloud endpoint, that saves not just money but often patience as well.

Reasoning and Logic: Correct Thinking, but Not Deep Enough

The Reasoning section is probably the most revealing part of this review, because it exposes the model’s core contradiction. Hermes 4 70B solves tasks correctly often enough, but without the didactic depth one would expect from a reasoning-oriented model in this size class. In the metacognition test with the two guards, for instance, the basic solution was clean: the classic question about what the other guard would say, followed by choosing the opposite door. The format with <thought> tags was also formally correct. But the substance was thin.

The Judge rightly criticizes not the logic itself but the elaboration. Hermes 4 70B explains the double-negation of the lie only loosely, offers no genuine exploration of alternatives, and forgoes any visual or tabular aids. The result is an answer that is correct but pedagogically sparse. For a quick assistant, that’s acceptable. For a 70B model classified as reasoning-capable, it shows insufficient ambition.

The pattern recurs in the module’s overall score: 66.42 percent in Logical Reasoning. That’s not weak, but it’s also not a performance that makes the label “Deep Thinking” convincing. The finding aligns with the architecture category: in default mode without Extended Thinking activated, Hermes 4 70B behaves more like a disciplined Instruct model with solid analytical reflexes than like a model that takes genuine pleasure in pulling a problem apart. Anyone expecting more depth will very likely want to activate Thinking mode deliberately in production. The benchmark intentionally did not do that. What was evaluated was the factory setting, not the optional upgrade.

Code Quality and Security: Capable, but Not Forensically Sharp

In the Code Quality Audit module, Hermes 4 70B reaches 62.96 percent. That’s the zone where a model can be useful without qualifying as a security auditor. The qualitative security test makes this concrete: Hermes identified 15 of 19 expected vulnerabilities in a PHP application, delivered the response in German, maintained the table format, and categorized many findings sensibly by severity.

The problem lies in the gaps. Among the missed items were some of the most classic, hard-hitting findings: hardcoded API secret, hardcoded database credentials, missing CSRF protection, and a reset token without expiration. These are not exotic edge cases — they’re standard entries from the OWASP toolkit. When a 70B model in security mode overlooks four central issues like these, it’s no longer a matter of stylistic preference. It’s a precision failure.

There are also misweightings. Type juggling was rated too mildly; session fixation likewise. Hermes 4 70B lists vulnerabilities neatly but often explains them only at surface level. The dangerous chaining of weaknesses — the real attack path from a seemingly minor individual finding to full system compromise — remains underexplored. That’s exactly where a security analysis separates itself from an enhanced checklist.

The model is not a fraud. It finds much of what’s relevant, maintains formats, and stays readable. But it doesn’t approach security with the cold precision of a penetration tester — more with the reasonableness of a tidy review assistant. For first-pass audits, that’s workable. For sign-offs in critical environments, it’s not sufficient without human follow-up.

CLI and Tool Proximity: Strong Working Model with a Dangerous Imagination

The CLI benchmark comes in strong at 87.34 percent. That fits the operational character of Hermes 4 70B well. When tasks are clearly defined and require precise output, the model frequently delivers exactly the form expected. The Tool Execution score of 88.33 percent also reads well at first glance. Here Hermes demonstrates why it positions itself as a fast assistant for developer and automation environments.

But this area has a catch, and it’s not cosmetic. In the Tool Use section, two hallucination findings occurred. In two tasks, Hermes 4 70B generated content that did not originate from the actual tool result retrieved but was fabricated. The P2 score was therefore capped by a hallucination ceiling in both cases. For content-critical tasks such as research, incident reports, or fact-bound summaries, this is a disqualifying signal. A model that doesn’t just interpret tool output but partially supplements it with invented content operates like an intern with a polished manner and poor source discipline. That is not a charming trait in a company’s engine room.

The implication is clear: Hermes 4 70B works well for tool-assisted tasks when a human has the final word or results are rigorously validated. For fully automated agent chains where tool output is passed as a reliable source into the next step, this tendency toward invention must be intercepted. Otherwise, efficiency very quickly becomes error acceleration.

Content Transformation: Creative, Structured, but the Word Limit Slips First

At 74.17 percent, Content Transformation is one of the model’s stronger disciplines. Hermes 4 70B can reshape texts, structure scripts, adjust tone, and generally remains readable and controlled. The qualitative video script test even shows genuine talent: the response was fully in German, included both analytical and transformational elements, and incorporated timestamps, screen notes, B-roll, music cues, and an Easter egg. That’s not raw text generation — it’s recognizably craft-trained assistance.

Yet the model fails on precisely the one constraint that is least negotiable in editorial and production contexts: length. In one task in the Content Transformation section, Hermes 4 70B exceeded the explicit word limit of 900 words, delivering 1,225 words136 percent of the limit. The system applied an automatic deduction of 18.00 points, corresponding to 20 percent of the achievable partial score. The substantive quality of the response is irrelevant at that point. The penalty applies regardless.

The length problem is not an isolated cosmetic flaw — it’s a genuine production deficiency. The Judge describes the script draft as engaged and formally complete but unusable in terms of timing. A requested three-to-five-minute video effectively became a significantly longer monologue. This is the moment where Instruct compliance and creative expansion work against each other. Hermes wants to deliver, and then delivers too much. For marketing and adaptation work, that’s often fixable. For strictly timed formats, it simply costs rework.

Documentation Quality: The Strongest Writing Module with an Unnecessary Own Goal

In Documentation Quality, Hermes 4 70B scores 74.63 percent and shows one of its more pleasant sides. The model writes in a structured, matter-of-fact style with relatively little filler. Especially for explanatory texts, that’s an advantage. It doesn’t tend toward ornamental fog but toward useful compression. In day-to-day work, that’s worth more than many ostensibly more creative styles.

However, this section also contains a hard rule violation. In one task in the Documentation Quality section, the model ignored an explicit language instruction and responded in English despite German being required. The extracted language comparison was unambiguous: DE=6, EN=29. That’s not a matter of taste or a sentence-level misunderstanding — it’s a clear instruction-following failure. In environments with a fixed target market, corporate wording requirements, or regulatory language mandates, a slip like this can lead directly to rejection.

This is particularly frustrating because Hermes 4 70B otherwise presents as well-steerable. The own goal doesn’t feel like structural incapacity — more like a loss of control under multiple simultaneous requirements. Regardless: anyone automating documentation generation with German as the target language must add a language check downstream. Otherwise, the wrong version will eventually reach the approval process, and no benchmark score will help at that point.

UX Writing and Cultural Intelligence: Polite, Capable, Rarely Brilliant

In UX Writing, Hermes 4 70B sits at 68.67 percent. That’s solid, but not a score that replaces a microcopy specialist. The tone is characteristic: correct, friendly, professional — but occasionally generic. The quality is reminiscent of a team member who has read all the style guides but hasn’t yet developed a genuine feel for product voice. That’s better than careless arbitrariness, but it rarely leaves the impression that someone is guiding the user with real precision.

Cultural Intelligence comes in at 68.92 percent and paints a similar picture. The qualitative excerpt on inclusive reformulation was largely successful: toxicity removed, gender-neutral phrasing applied, professional enough. But the solution sounded more like standard compliance than linguistic finesse. The Judge specifically criticized the choice of “Persönlichkeit” over a more idiomatically precise term like “Fachkraft,” and the loss of a more energetic tone. That hits the core: Hermes can handle culturally sensitive tasks, but the final nuance occasionally eludes it. It knows where not to cause offense. It less often knows how to win linguistically in the process.

Data Privacy and Data Sovereignty

The picture is mixed, and precisely for that reason it deserves a sober reading. The calculated Sovereign Risk is MEDIUM, grounded in a medium weights provenance risk. The developer is Nous Research Inc., headquartered in San Francisco, California, USA. Regarding the model source itself, the Vendor Card states: based on available data, Nous does not operate its own public API, provides open weights, and lists 0 days of data retention. At the same time, no GDPR DPA is available. For organizations required to procure in GDPR compliance, that’s not a footnote — it’s a potential compliance obstacle.

The reviewed Vendor Card describes the data location only as “Local or third-party hoster.” For this benchmark setup, however, the critical point is that Hermes 4 70B ran as a Cloud Open Weights model via a third-party endpoint. This shifts the data privacy question from the model source to the specific deployment infrastructure. No verified provider data for the deployment infrastructure was available here. For European organizations, this means in practice: the weights are open, but sovereignty depends entirely on the chosen hoster and its jurisdiction. As soon as a US provider or US-controlled infrastructure is involved, the CLOUD Act may become relevant. Physical processing in Europe does not automatically eliminate this risk.

Conclusion

Hermes 4 70B is an interesting model with a clear personality: fast, token-efficient, pleasantly steerable, and operationally useful. It reaches 70.68 percent, delivers 84.61 tokens per second, remains stable across 43 of 43 tests, and shows genuine practical relevance particularly in CLI, documentation, and structured content work. But it’s not a model that fully cashes in on its size in every core area. In Reasoning, it often stays correct without going deep. In Security, it’s helpful but not thorough enough for hard reviews. And in tool-bound work, the documented hallucinations are the point where trust must become caution.

For deployment, this means: Hermes 4 70B fits well as a fast working assistant for editorial teams, dev workflows, technical documentation, and format-constrained everyday tasks with human final review. It’s less suited for fully automated fact-critical agent chains, security sign-offs without review, and any setting where word limits, language requirements, and source fidelity are absolutely non-negotiable. This model is not a fraud, but it’s not a bulldozer either. It’s more of a capable operator with a slight tendency toward overconfidence. As long as you know that, there’s a lot you can do with it.

This evaluation was generated automatically based on the benchmark data. Model used: GPT 4.5 by OpenAI. The raw data and full methodology are documented in the GitHub project.