LLM Model Review
Created on · 36B · NVFP4 · Compressed-Tensors · 512K-Context · Long Context · Agentic Orchestrator
With an overall score of 67.14% and the Speed Profile Badge Batch Tool Expert, Hermes 4.3 36B (vLLM, Seed-OSS, Dense, NVFP4) is not a nimble chat all-rounder but a heavy-duty work model with planning ambitions and noticeable sluggishness. The editorial classification fits: as a Generalist in the Server class with Dense architecture, it must be measured against broad competence, not a specialist talent. That is precisely where it delivers a mixed picture — structured, often usable, occasionally clever, but too frequently imprecise in language, formatting discipline, and fact-critical tool use. Sovereign Risk: MEDIUM — Open Weights under Apache-2.0, operable locally without cloud egress; provenance nonetheless remains less than fully neutral due to US jurisdiction and downstream NVFP4 quantization.
Header Grades: Stability and Reliability
| Metric | Value | Rating | Analysis |
|---|---|---|---|
| Timeout Rate | 8/49 | Unreliable | The model is unreliable and drops out significantly often in practice. |
| P95 Response Time | 191.16 s | Critical | Extreme tail latency. The model scatters massively and is unsuitable for time-critical processes. |
Architecture and Operating Mode
The preliminary classification captures the character surprisingly well, if read correctly. Hermes 4.3 36B (vLLM, Seed-OSS, Dense, NVFP4) is a Dense model with 36 billion parameters — full capacity per token rather than expert selection as in MoE systems. This explains part of its behavior: more raw model mass per response, but also less agility. In the Server class, that is legitimate. In this class, however, it is also mandatory to be not merely presentable but load-bearing.
The actual test mode matters: this run took place in Standard Mode, meaning Thinking was disabled. This is not a misconfiguration but the fair out-of-the-box operating condition. Because Hermes fundamentally supports Thinking-Optional and is simultaneously classified as an Agentic Orchestrator, one may reasonably infer more internal planning than a plain instruct model would show. One may not, however, pretend that every slow or broadly unrolling output is automatically a sign of deeper intelligence. In the benchmark, what counts is what actually lands on the table.
The fact that the model can additionally be read as an Uncensored-Finetuned family is relevant as a leniency rule, but not a free pass. For Coding and DevOps, some tolerance for nuance compared to strictly production-optimized base models is reasonable. For hallucinations in tool workflows, that leniency does not apply. There, the posture ends and the risk begins.
Speed and Runtime Profile
Hermes carries the badge Batch Tool Expert. That is a fairly apt short diagnosis. This model is more plausible for longer, batch-style tasks than for direct, snappy interaction. Its speed is qualitatively low to moderate, and the variance weighs more heavily than the raw mean. Anyone looking for a local model that feels like a responsive console is in the wrong place. Anyone willing to accept longer responses, tool schemas, and more complex structural work gets something closer to a deliberate project processor.
The local model was tested natively on an ASUS GX10 / NVIDIA DGX Spark (GB10 Grace Blackwell Superchip, ~115 GB Unified Memory — no practical memory limit for tested model sizes). For the assessment, what matters most is therefore this: the observed outliers should be taken seriously as a real stability and latency finding, not dismissed as a trivial artifact of a constrained setup.
On the positive side, token economy is at least solid. Hermes behaves token-economically; no module exceeds the expected verbosity envelope. The model is not slow because it rambles endlessly. It is simply not an express train.
Reasoning and Logic
For a model with Thinking lineage in its family tree, the reasoning domain is respectable but not intimidating. The logic score sits at 62.23%. That is not a collapse, but for a 36B Dense model with reasoning ambitions it is also not cause to reach for the trumpet. In stronger cases, Hermes shows clean basic structure, traceable steps, and a certain didactic order. In the guard riddle, for instance, it lands on the correct core question and explains the double inversion accurately. That is not trivial — many models fail precisely there by tripping over their own explanation.
What is missing is the second layer: clarity under pressure, pedagogical sharpness, elegant compression. The Judge protocol rightly notes that the presentation remains functional but never becomes excellent. Visual or tabular aids are absent, alternative formulations are missing, and occasionally so is the final precision in the why. The model does not reason incorrectly. It simply often does not reason all the way through — at least not visibly in the result.
In one reasoning task, the model also ignored an explicit language instruction and responded in English. This is not a technical glitch but a clear weakness in instruction-following. In environments with a fixed target language, exactly this kind of slip is not a minor offense but a ticket for manual review.
Code Quality and Security
The Code Quality domain is a good example of why this model should be neither written off prematurely nor praised prematurely. The module score of 66.8% is decent but not rock-solid. In security analysis, Hermes correctly identifies many classic vulnerabilities: SQL Injection, plaintext passwords, XSS, session issues, Path Traversal, CSRF. That is the solid craft baseline one should expect from a large generalist.
Then come the cracks. In a security audit, the model responded entirely in English despite an explicit requirement for German. This costs not only presentation points — it points to an old problem with many large models: once structure, technical content, and table formatting converge, the language requirement is the first thing to fall off the wagon. Substantively more serious is the fact that several central vulnerabilities are missing, including IDOR, hardcoded secrets, and a missing expiry date for reset tokens. IDOR in particular is not a detail in such chains — it is often the lever that turns a sloppy application into a full account takeover. When a model sees the fog but does not name the abyss, that is not an academic aesthetic flaw.
From a security standpoint, this means: Hermes is suitable for initial triage, not for final sign-off. It catches a lot. It does not always prioritize correctly. And in security, half-knowledge is not half as good as full knowledge — it is twice as dangerous.
Tool Use and Hallucinations
This is the Achilles’ heel of this run. The Tool Use score of 43.33% is weak, and the logs state the reason unambiguously: in four tool tasks, the model hallucinated content that did not originate from the retrieved tool result. The affected tasks were tooluse002, tooluse004, tooluse005, and tooluse006. The system applied the hallucination cap to the P2 score in each case. For research, fact-critical summaries, and agentic reports, this is a disqualifying signal.
Precisely because Hermes is architecturally classified as Tool-Use-capable and as an orchestrating model, this finding is serious. An orchestrator may receive some leniency on exact one-liners. It may not invent the contents of its own tools. Anyone who treats tool results as mere rough inspiration is not building an agent — they are building a self-confident error source.
This is the point at which theoretical agentic competence becomes practical distrust. As long as every external source is cross-checked, the model can be used. Without that safeguard, it should not be deployed as a factual endpoint.
UX Writing and Content Transformation
In UX Writing, Hermes displays a pattern typical of this model: usable practice, limited depth. The module score of 60.85% is not explained by gross failures but by missing psychological precision. In a skip-button optimization task, the model gets many things right: clearer language, less jargon, better pacing, more concrete phrasing. What is absent are the underlying behavioral principles, quantitative targets, and methodological rigor that elevate decent copy to convincing UX work. Add to that scope creep: an unsolicited supplementary block on psychological principles that dilutes focus rather than strengthening the response. Useful, yes. Disciplined, no.
In the Content Transformation domain, the picture is harder. The score of 69.42% looks acceptable at first glance, but the logs reveal a recurring weakness when requirements combine language, length, and format. The model can imitate structure but drops individual core conditions under load. This is clearest in a video script task: timestamps, production notes, and a casual tone are formally present, but the primary language shifts to English even though German was required. That is not a minor slip — it is a direct failure to meet a mandatory condition.
In one Content Transformation task, the model exceeded the explicit word limit of 250 words by 210%, producing 525 words. The system applied an automatic deduction of 33.60 points, equivalent to 40% of the achieved partial score. The substantive quality of the response is therefore irrelevant. The penalty applies regardless.
This violation is worth noting because it reveals the model’s character. Hermes has ideas, structural instinct, and occasional stylistic sensibility. But when multiple guardrails apply simultaneously, it apparently treats the word limit as a non-binding suggestion. For editorial or agentic workflows, that is unwelcome. A model that writes cleanly but overruns briefs creates more rework than it saves.
Documentation Quality
With 62.58% in Documentation Quality, Hermes stays in solid territory without distinguishing itself. Its strength lies in usable structure and a certain willingness to bring the reader along. The weakness is an absence of sharpness in prioritization and compression. The model tends to explain a touch more broadly than ideal for precise technical documentation, without that breadth translating into genuine completeness.
This fits the general character of this run. Hermes is not terse, but not verbose either. Not confused, but often one step away from clear excellence. For internal documentation drafts, that is manageable. For texts that must hold up in critical environments without post-processing, less so.
Cultural Intelligence
The 74.0% in the Cultural Intelligence module ranks among the more encouraging scores for this model. Hermes generally works cleanly in terms of language and stays in German when it matters. At the same time, this domain also shows that good intent and good execution do not always coincide. In the revision of a problematic job posting, the model does remove part of the toxic and aggressive tone, but fails precisely at the core requirement: reliably neutralizing gender bias. Choosing “Fachmann” instead of an inclusive form such as “Fachkraft” is not a peripheral error — it is a miss at the central briefing point.
There is also a tone problem: negative formulations are in some cases merely relabeled rather than genuinely rebuilt. Professional, inviting language does not emerge from that. The result is better than the source text, but it has that slightly mechanical awkwardness one immediately recognizes when the actual goal is reaching real applicants. Not wrong enough to fail. Not refined enough to trust.
CLI and Operational Precision
The CLI score of 82.22% is one of the clear strengths of this run. That aligns surprisingly well with the Orchestrator classification. In the operational domain, Hermes frequently delivers the necessary structure, recognizes task patterns, and moves through shell and command logic reliably enough to be useful in assistant roles. That it looks better here than in Tool Use is not a contradiction. CLI tasks reward structure and operational soundness. Tool Use penalizes invented facts. Hermes handles one better than the other.
For semi-automated technical workflows, this is an interesting combination: good operational orientation, but only under the condition that external results are not blindly accepted. Anyone using Hermes to draft commands or structure work plans can benefit. Anyone treating it as a truth machine over tool outputs is building on sand.
Data Protection and Data Sovereignty
A dedicated data protection alert block is not necessary here, since this is not a cloud endpoint but Open Weights in local operation. The provenance is nonetheless relevant: the calculated Sovereign Risk is MEDIUM. The reasons are the US origin of Nous Research, the associated CLOUD Act jurisdiction on the provider side, and the downstream NVFP4 quantization of the publicly available base weights. For European organizations, the practical good news is decisive: in fully local operation, no prompts or content are transmitted to any provider. The risk therefore lies more in the origin and governance of the weights than in ongoing data egress.
Conclusion
Hermes 4.3 36B (vLLM, Seed-OSS, Dense, NVFP4) is a model with substance, but without the effortlessness one would like to see in its class. It reaches 67.14%, demonstrating enough capability to be taken seriously, but too many weaknesses to be trusted blindly. As a Generalist in the Server class with Dense architecture, it is entitled to work broadly. That is exactly what it does. It writes usably, plans adequately, delivers solid CLI work, and produces reasonable baseline logic. Then it stumbles over language requirements, ignores word limits, and hallucinates in tool tasks precisely where factual fidelity is non-negotiable.
The comparison to the second run of the same model is narrow but instructive: the Standard Mode tested here, at 67.14%, sits slightly below the separately available Thinking variant at 68.02%. The gap is small, but the character shifts. Thinking yields somewhat more structural gain, yet does not transform Hermes into a new type. It remains a deliberate tool-carrier rather than a brilliant problem-solver even then.
For practical deployment, this means: sensible for local assistance, documentation drafts, CLI support, and longer structured tasks with human oversight. Not sensible as an unsupervised research agent or as a reliable final authority for security or fact-critical workflows. The open Apache-2.0 license, local operability, and large native context are genuine advantages. But character does not substitute for reliability. And reliability is precisely what this Hermes most conspicuously lacks.
This evaluation was generated automatically based on the benchmark data. Model used: GPT 4.5 by OpenAI. The raw data and full methodology are documented in the GitHub project.