LLM Model Review
· Instruction-Tuned
With an overall score of 68.99 percent, Hermes 4 70B makes very clear what it wants to be and what it does not: a fast, direct cloud Open Weights model via OpenRouter, operating in standard mode without explicitly activated Extended Thinking, that prefers decisive answers over performative deliberation. The Speed Profile Badge reads “Real-Time Tool Expert.” In practice, that means: clearly optimized for quick, interactive tool and workflow tasks, not for the long literary breath. For a dense 70B server model with reasoning ambitions, that is respectable — but also contradictory: plenty of speed, too little depth at precisely the points where a reasoning-oriented model should deliver. Sovereign Risk: MEDIUM — Hermes comes from Nous Research in the USA; as Open Weights, the source risk resides in the model itself, and when used via third-party cloud providers, jurisdiction and potential government access remain a real compliance concern.
Header Notes: Stability and Reliability
| Metric | Value | Rating | Analysis |
|---|---|---|---|
| Timeout Rate | 1/49 | Sporadic | The model shows sporadic dropouts that would require retries in practice. For a cloud Open Weights endpoint, this is not a puzzle of model logic but a concrete reliability risk at the API or network level. |
| P95 Response Time | 20.54 s | Consistent | Very low tail, almost no outliers. For a server candidate of this size, that is a pleasantly short shadow of latency. |
Architecture and Character: Quickly Instructed, Only Limitedly Inclined to Think
The metadata “Instruct, Thinking-Optional” is no minor detail here — it is the key to the model’s character. Hermes 4 70B is a dense 70B server-class model, optimized for direct instruction execution and structured responses. At the same time, the model family signals reasoning ambitions. However, what was tested was the standard mode with n/a, meaning no separate thinking switch. That is precisely why the results must be read carefully: shorter, more direct answers in this run are not misconfiguration — they are the normal operating state.
The problem is simply that Hermes 4 70B does not always find the right balance, even given this calibration. It is often fast and clean in format, but it occasionally economizes in the wrong places. Especially for a reasoning-oriented use-case profile, one does not necessarily expect epic length, but one does expect visible argumentative substance. Instead, the model at times resembles a good editor working under too tight a deadline: correct, usable, but leaving the crucial connecting sentences in their head.
Performance: Speed as a Strength, Not an Excuse
The “Real-Time Tool Expert” badge fits surprisingly well. Hermes 4 70B generates output very quickly for a model of this class. This speed should be read explicitly as a performance value of the cloud provider — in this case the OpenRouter endpoint and its underlying cloud infrastructure. Such values are not an abstract natural law of the model, but the result of performant server delivery plus network path. For users, this is still relevant, because that is exactly what the service feels like day to day.
The flip side of this speed is more subtle. Hermes 4 70B does not behave like a classic heavy thinker, but like a tool model with a reasoning facade. That is not inherently bad. In CLI and tool environments, direct responsiveness can be more valuable than majestic thoroughness. But as soon as tasks demand multi-step explanation, didactic grounding, or precise creative completeness, it becomes apparent that the engine prefers to set off rather than fully unfold the map.
Reasoning and Logic: Right Thinking, But Not Deep Enough
At its logical core, Hermes 4 70B is better than its overall impression suggests. On the classic guard puzzle, it arrives at the correct solution and formulates the central counter-question accurately. That is the good news. The less good news: the explanation stays below its potential. The Judge does not fault the final answer, but the missing proof. The model does not cleanly demonstrate the mechanics for both cases, only hints at the double negation, and even commits an unfortunate ambiguity with the phrase “door to certainty,” where clarity about the freedom door versus the death door was needed.
That is more than a cosmetic flaw. A reasoning-oriented server model must not only be right — ideally it must bring the reader along in a way that makes the solution path feel robust. Hermes 4 70B comes across here more like a student passing an exam who has the right result but only implicitly supplies the essential intermediate steps. You can grant it the answer. You just cannot trust it blindly.
At least the model does not refuse the required <thought> tags in the metacognition protocol at hand — it uses them. No systematic compliance failure is detectable on the basis of the available excerpts. That matters, because it shows: the problem here is not stubborn policy, but limited argumentative elaboration.
Code Quality and Security: Tidy on the Surface, Porous in Depth
The code quality module reveals the great ambivalence of Hermes 4 70B. Formally, it works cleanly. The Markdown table is correct, the response stays in German, the structure is maintained. That sounds trivial, but it is not. Many models fail at exactly this unspectacular discipline. Hermes 4 70B fails elsewhere: in completeness.
In the evaluated security audit, the model identifies only 9 of 19 relevant vulnerabilities. That is not a narrow defeat — it is a coverage gap of 53 percent. What is particularly critical is not that some exotic detail is missing, but that entire central categories are absent: IDOR, XSS, weak reset token generation, missing CSRF protection, hardcoded secrets, problematic header forwarding, missing token expiry. Anyone who wants to conduct a serious security review does not need a model that spots half the fires and misreads the rest as ambient temperature.
The quality of the findings is also mixed. SQL injection in the login is identified, but without a concrete attack illustration. Session fixation is mentioned, but not really dissected. Path traversal appears only as a generic “insecure file access,” not as a precisely described escape path. The model is not blind. It simply does not look closely enough.
For a model with Hermes genetics and a more open alignment, some leniency is appropriate — but not absolution. As an uncensored-finetuned relative, its focus leans more toward steerability and low Refusal rates than toward pedantic security analysis. Even so: at server class and 70 billion dense parameters, this security coverage is simply too thin. The model delivers a usable first pass, but not an audit on which one should base a sign-off.
Content Transformation and UX Writing: Talent Present, Discipline Missing
Hermes 4 70B can write. That is not the question. The question is whether it also finishes writing under real constraints. In the content transformation module, it initially demonstrates exactly the kind of competence one wants from a well-trained instruct model. The video script on two-factor authentication is clean German, with timestamps, production notes, screen annotations, and a natural spoken-word tone. The model understands what production-ready output looks like. It even builds in an Easter egg and stays comfortably within the available budget. What is missing is not formal awareness, but emotional precision. The hook is too tame, the dramaturgy around backup codes too flat, the timing calculated a little too tightly. Solid craft, no direction with goosebumps.
In another task within the same module, however, Hermes 4 70B responded in English even though German was explicitly required. That is not a peripheral error — it is an automatic Hard-Constraint violation due to language mismatch. The substantive quality of the response becomes secondary. When the target language is firmly specified, the model simply fails in practice. Together with the language error in the documentation section, a pattern emerges: under simultaneous constraints of language, length, and format, Hermes 4 70B drops the language requirement as the first condition. That is no longer an outlier — it is a structural weakness in instruction following.
In UX writing, it gets harder. There, Hermes 4 70B starts strong, analyzes precisely, and formulates psychologically informed optimization suggestions. But only a fraction of the actual assignment was visibly delivered. The rest ends in a technical abort. That is precisely what makes the failure so frustrating: the model is clearly capable of the task, but manages its resources poorly. It thinks and formulates until it runs out of steam.
In the UX writing section, one output breaks off mid-table. The response is technically aborted, not a content error. The score deduction results from the incomplete response, not from substantive shortcomings.
This episode is more relevant to everyday use than any stylistic fine-tuning. In agent chains, review pipelines, or editorial workflows, what ultimately counts is not whether a model had good ideas, but whether the complete payload arrives. Hermes 4 70B displays the uncomfortable combination of intelligence and missing self-discipline. A strong first third does not rescue two missing thirds.
Documentation: Competent, But Not Language-Faithful Enough
The documentation scores are overall only mediocre, and the qualitative evidence points to a familiar problem: Hermes 4 70B can explain things in a structured way, but loses the last degree of precision under multiple simultaneous constraints. Particularly serious is the documented language violation in a documentation task where German was required but the model responded in English. That too is an automatic Hard-Constraint violation due to language mismatch. For organizations with a fixed documentation language, this is not a cosmetic error — it is an immediate exclusion criterion for uncontrolled use.
In documentation workflows especially, language consistency is not a luxury but process hygiene. If a model holds to German in security audits but spontaneously switches to English in documentation, that is not a stylistic quirk. It is unreliability in a core requirement.
Cultural Intelligence: Pleasingly Clean, But Not Fully Idiomatic
In the Cultural Intelligence module, Hermes 4 70B shows one of its more appealing sides. The task of inclusively revising a problematic job posting is handled entirely in German, without explanatory ballast and with a professional tone. Problematic terms are removed, gender bias is largely eliminated, and output discipline is sound. That is no revolution, but good practice.
The Judge rightly notes that the phrasing does not always hit the idiomatic ideal of German HR communication. “Dynamische Persönlichkeit” sounds somewhat more marketing-heavy than necessary, and “Fachmann” remains weaker than a truly neutral choice such as “Fachkraft.” But these are nuances at a good level. The core finding matters: Hermes 4 70B understands social language hygiene and does not tip over into wooden moral prose in doing so. That is worth more than some benchmarks suggest.
Tool Use and Hallucinations: Strong on Access, Risky on Output
The tool execution value comes out noticeably better than the actual tool use score. That fits the model’s character. Hermes 4 70B can apparently handle tool context, but stumbles at the synthesis of tool result and final response. That is also precisely where the most serious hallucination finding sits.
In one tool use task, the model generated content that did not originate from the retrieved tool result but was fabricated. The score was therefore capped via hallucination cap. For content-critical applications such as research, factual reports, or compliance-adjacent summaries, this is a disqualifying signal. A tool model that improvises again after looking something up is like an intern who requests files and then quotes from memory.
This weakness stands in striking contrast to the speed badge. “Real-Time Tool Expert” describes the interaction aptly. When it comes to the reliability of extracted content, the title should be read with a footnote.
API Cost Profile
Hermes 4 70B is a cloud Open Weights model, so not only quality but also token discipline matters as a cost signal. In almost all modules, the model behaves pleasingly economically. There is one exception, however: in UX writing, it produces an average of 3,518 tokens against a fleet median of 1,644. That corresponds to a factor of 2.14 relative to the average across all tested models.
That would be easier to accept if quality were simultaneously dominant. It is not. On the contrary: it is precisely in this module that one task ends in an abort. For API usage, this means plainly: higher costs, longer responses, no guaranteed quality gain. Hermes 4 70B is not merely verbose there. It is inefficient.
Data Privacy and Data Sovereignty
The data privacy situation for Hermes 4 70B is two-sided. The model itself comes from Nous Research Inc. in San Francisco and is provided as an Open Weights model under a Modified MIT license. According to the Vendor Card, Nous does not operate its own public API, based on the reviewed company website. The actual data privacy risks therefore only arise through the chosen deployment path — in this case, concretely through use via OpenRouter as a cloud proxy.
The calculated Sovereign Risk is MEDIUM. This is justified by a medium-level weights provenance risk. A GDPR DPA is not available according to the Card. For organizations that must operate in GDPR compliance, this is not a detail but a potential compliance obstacle. The Card lists the data location as “Local or third-party hoster” — meaning no reliably fixed processing location. Data retention is stated as 0 days. That sounds good, but does not substitute for contractual assurance. For German and European users, the conclusion is therefore: the open model is more sovereign than a closed black box, but the concrete data privacy reality depends on the host. And as soon as a US provider or US-connected infrastructure is involved, the CLOUD Act remains legally relevant, even if data is physically located in Europe.
Conclusion
Hermes 4 70B is a model with a clearly recognizable temperament. It is fast, structured, often pleasantly direct, and better at culturally sensitive language tasks than one might hastily assume of an openly aligned 70B model. At 68.99 percent, it delivers an overall serviceable benchmark performance. But precisely because it is a dense server model with reasoning-oriented ambitions, its deficits stand out more sharply. Security analyses remain too shallow, UX outputs can abort technically, language requirements break down across multiple modules, and in the tool context a documented hallucination is not a cosmetic flaw but a warning sign.
For interactive tool tasks, general reformulations, initial analyses, and structured working drafts, Hermes 4 70B is genuinely interesting — especially since its speed via OpenRouter is impressive for a model of this class. For security-critical audits, fact-critical research synthesis, or unsupervised agent workflows, however, it lacks the final degree of reliability. Hermes 4 70B is not a bluffer. But it is a model that too often behaves as though it considers the first usable draft the final version. In editorial terms, one would call that talented — but not yet ready to air.
This evaluation was generated automatically based on the benchmark data. Model used: GPT-5.4 by OpenAI. The raw data and full methodology are documented in the GitHub project.