LLM Model Review
Created on · 36B · NVFP4 · Compressed-Tensors · 512K-Context · Long Context · Agentic Orchestrator
With an overall score of 68.02%, Hermes 4.3 36B (vLLM, Seed-OSS, Dense, NVFP4) presents a peculiar profile: as a generalist in the Server class with 36 billion Dense parameters, it aims to think, orchestrate, and use tools simultaneously — yet in the benchmark it remains a model with a noticeable streak of stubbornness. The concrete test run was conducted in Thinking mode, and it shows: responses are often structured and deliberate, but not automatically more precise. The Speed Profile Badge is Batch Tool Expert; that stands for a rather unhurried, tool-oriented usage pattern rather than interactive back-and-forth. Sovereign Risk: MEDIUM — the Open Weights originate from a US-shaped legal and provenance context; when run locally, no prompts flow to a cloud provider, but the origin remains a relevant compliance factor nonetheless.
Header Notes: Stability and Reliability
| Metric | Value | Rating | Analysis |
|---|---|---|---|
| Timeout Rate | 7/49 | Unreliable | The model is unreliable and drops out at a significantly high rate in practice. |
| P95 Response Time | 160.89 s | Critical | Extreme tail latency. The model exhibits massive variance and is unsuitable for time-critical processes. |
Architecture and Character
The pre-assigned metadata classification captures the essence surprisingly well — just not quite as flatteringly as the tag chain might suggest. Hermes 4.3 36B (vLLM, Seed-OSS, Dense, NVFP4) is a generalist with Server-class ambitions, Dense architecture, and a long context window of 512K tokens. Dense here means not merely the technical “classic transformer,” but practically also: all 36 billion parameters are active for every token. You pay for capacity not only at load time, but also at runtime.
The dual nature of the categorization matters. Architecturally, the model is marked as both Thinking and Thinking-Optional. For this report, the hard runtime mode is what counts: testing was conducted explicitly with Thinking enabled. That is not a side note — it is the evaluation framework. Longer, argumentatively developed responses are not a flaw here, but an intention. For a model with Agentic Orchestrator ambitions, one should therefore expect it to decompose tasks cleanly, recognize priorities, and not stumble at the first formality. That is exactly the standard against which Hermes is measured. And exactly where it chafes in the benchmark.
Speed and Efficiency
Hermes is a local model and was evaluated natively on ASUS GX10 / NVIDIA DGX Spark (GB10 Grace Blackwell Superchip, ~115 GB Unified Memory — no practical memory limit for tested model sizes). The Batch Tool Expert badge fits the observation: the model is not built for fluid dialogue, but rather for tasks where a longer processing time is acceptable as long as a usable working result emerges at the end. That can be legitimate in documentation or research contexts. For agents that iterate quickly or wait on responses in chains, it is a noticeable bottleneck.
At least Hermes behaves token-economically. No module exceeds the expected verbosity range. That is a genuine plus, because slow models that also ramble penalize patience twice over. Hermes does not do that. Across CLI, Code Quality, Documentation, UX, and Content it consistently produces fewer or roughly as many tokens as the fleet median. The slowness here does not come from textual excess, but from the character of the model itself. That is more honest, but no more comfortable.
Reasoning and Logic
In the Logical Reasoning module, Hermes lands at 63.34 points. That is not catastrophic, but for a model running in Thinking mode that architecturally bills itself on reasoning, it is not a commanding performance either. The qualitative logs reveal a recurring pattern: the solution idea is often present, but the didactic or formal execution remains shallower than it should be.
Particularly telling is the logic task with the two guards. There, the actual inference is correct: Hermes understands the principle of double negation and arrives at the right door. The problem is not the thinking, but the packaging. The Judge criticizes an overly superficial explanation of the mechanism, missing alternative formulations, and less pedagogical depth than the reference standard. For everyday use, that is still acceptable. For a Thinking model in the Server tier, it is a small betrayal of its own ambitions.
There is also a concrete language error in the Reasoning section. In a metacognitive task, the model ignored the explicit language instruction and responded in English. That is not a cosmetic flaw — it is an instruction-following weakness. In environments with a fixed target language, this does not register as an intellectual misunderstanding but as an immediate production failure.
Overall, the reasoning is serviceable, but not authoritative. Hermes visibly engages in thought. It is just that thinking here does not automatically translate into a quality advantage. At times it feels more like an elaborate working path to a merely adequate result.
Code Quality and Security
With 67.8 points in the Code Quality module, Hermes delivers no embarrassment — but also no security report you would want to sign off on without review. The model correctly identifies many classic vulnerabilities, formats cleanly, and stays disciplined in the table. SQL injection, plaintext passwords, XSS, session fixation, path traversal: the repertoire is there. A solid foundation is clearly present.
But the weaknesses are uncomfortable precisely where a security context allows no shrugging. In one audit, Hermes missed 4 out of 19 relevant vulnerabilities — a gap of 21%. These were not decorative footnotes; they included hardcoded secrets, root credentials without a password, a reset token without an expiry time, and a header-context issue. More problematic still is the misclassification of individual findings. An IDOR case — unauthorized access to objects via manipulated identifiers — was logged as missing input sanitization. Those are not the same thing. Confusing threat models means ultimately confusing priorities as well.
The fix suggestions are also not always precise. In the token injection case, the model points to prepared statements or mysqli_real_escape_string, even though the actual vulnerability was in a mail context. That is not a minor slip — it is a wrong diagnosis. A security model that hears the right alarm but points to the wrong circuit is useful for a preliminary pass and risky for sign-off.
For context: Hermes also carries the character of an uncensored fine-tuned model. Such models are not primarily built for DevSecOps fidelity, but for freer content generation on top of an intact base architecture. That explains some things, but does not excuse everything. In practice, Hermes in a security context is better suited as an assistant for a first sweep than as a final authority.
Documentation, UX, and Content Transformation
The most charitable reading of Hermes in the writing domain is: functional, often structured, occasionally rough around the edges. The less charitable reading: it fulfills the obligation and regularly forgets that form is also part of the task.
In the Content Transformation area, both sides show up at once. On the positive side is the fundamental ability to restructure material, organize it, and adapt it to target media. The log for the video script shows a usable result with analysis, timestamps, production notes, and an Easter egg. That is not trivial. Many models fail simply at keeping the parts of the task together. Hermes does not.
Only here the devil is not in the details — it is directly in the production reality. The script claims a runtime of 4:15 minutes, but the visible amount of text suggests closer to 5:40 minutes. The timestamps do not add up cleanly to the actual speaking duration. For editorial teams and creator workflows, that is a blocker. A script with faulty timing logic is like a timetable with poetic license: well-intentioned, practically useless.
There is also a hard constraint violation in the Content module. In one task, the model exceeded the explicit word limit of 250 words with 419 words — 168% of the limit. The system automatically applied a deduction of 16.80 points, or 20% of the achievable score. The substantive quality of the response is therefore irrelevant. The penalty applies regardless. That is precisely what makes such violations so frustrating: it was not the Judge being pedantic, but the ruleset being consistent.
The language discipline issue weighs even heavier. In another Content task, Hermes ignored the explicit language instruction and responded in English. This is not an isolated slip, but part of a broader pattern across multiple modules. When faced with simultaneous constraints on language, length, and format, the model demonstrably drops the language requirement first.
In Documentation Quality, the score stands at 66.94 points — above the model’s overall impression, but not above its structural flaws. Here too, one task produced a language error: instead of the required German, Hermes delivered English. In production documentation pipelines, that is an immediate failure, because translation or editorial rework should not be absorbed “somehow” downstream. Documentation lives on reliability, not good intentions.
UX Writing comes in at a rather pale 62.05 points. That fits the overall character. Hermes can rewrite, smooth out, and neutralize toxic or exclusionary language. In the qualitatively logged example, it replaces problematic phrasing adequately and adheres to the instruction to output only the rewritten German text. What is missing is often warmth, idiomatic elegance, and the final sensitivity to tone. The Judge describes it aptly: functional, but less inviting and less natural than the reference. You get the result. It just sounds, in places, as though a correct model left its charm at the coat check.
The language failures are not isolated outliers. Across multiple tasks in the Content, Documentation, and Reasoning modules, the model shows a consistent pattern: when faced with simultaneous constraints on language, length, and format, it drops the language requirement as the first condition. Concretely, this affected a video script task in the Content Transformation module, a documentation task, and a metacognitive logic task. For German-language production environments, this is a structural risk, not a detail.
CLI, Tool Use, and Hallucinations
What is genuinely interesting about Hermes is its tool ambitions. The metadata promises Tool Use, Long Context, and Agentic Orchestrator capabilities. The benchmark result, however, reveals a clear gap between claim and deployment safety. The ToolUse score sits at 67.46, with Tool Execution itself at 90.0. Simplified, that means: the model can formally integrate tools and handle the basic flow. But when it comes to fidelity to the actual tool output, things get precarious.
Hallucinations
This is the model’s Achilles’ heel. In three Tool Use tasks, Hermes generated content that did not originate from the retrieved tool result but was fabricated. Each instance triggered a hallucination cap on the score. For content-critical tasks such as research, factual reports, or agentic summaries, this is not a minor cosmetic flaw — it is a disqualifying signal. A tool model must not pretend to have looked something up when it has in fact improvised.
Precisely because Hermes is architecturally classified as an Agentic Orchestrator, this finding carries extra weight. An orchestrating model does not need to execute every string perfectly. It may occasionally benefit from specialized sub-agents for shell subtleties or exact one-liners. What it cannot afford is freely embellishing tool outputs. Because that is exactly where orchestration ends and fiction begins.
The comparison with the standard mode of the same model is soberly interesting. The Thinking run improves the overall score slightly from 67.14% to 68.02% and lifts Documentation and ToolUse somewhat. At the same time, Content Transformation, Cultural Intelligence, and UX Writing dip slightly. Thinking does not make Hermes categorically better. It makes the model somewhat more analytical and somewhat less smooth. Depending on the task, that can be a gain or a source of friction.
Cultural Intelligence
With 67.3 points in the Cultural Intelligence module, Hermes falls short of its better moments. That is a shame, because the qualitative probe in particular shows that the model can competently defuse problematic terms, toxic metaphors, and gender-coded phrasing in German. It clears away the coarse disturbances without moralizing unnecessarily. That is practical and often exactly right.
What is missing is fine motor control. The Judge criticizes clunky formulations such as “geselligem Abenteuern” and an overall more neutral, less inviting tone than the reference text. Hermes removes the toxins, but does not cook them into a good dish. For internal rewrites, that is often sufficient. For public-facing communication that demands idiomatic precision, less so.
Privacy and Data Sovereignty
A dedicated cloud privacy section is not necessary here, since Hermes 4.3 36B (vLLM, Seed-OSS, Dense, NVFP4) is run locally as an Open Weights model. What remains relevant is provenance: the Weights Provenance Risk is rated MEDIUM, because the publicly available Apache 2.0 weights originate from a US-shaped context and Nous Research, as a US company, is subject to the CLOUD Act. For local deployment, this substantially reduces the operational data exposure, but does not resolve every governance question around origin and compliance documentation.
Conclusion
Hermes 4.3 36B (vLLM, Seed-OSS, Dense, NVFP4) is not a bad model. But it is one you need to understand before you trust it. As a local generalist in the Server class with 36 billion Dense parameters, it brings a genuine toolkit: long contexts, Thinking mode, Open Weights, Apache 2.0 license, serviceable analytical capabilities. At its best, it feels like a deliberate technical writer with a grounding in security and a tendency toward over-explanation. At its worst, like that same writer who misreads the brief at three critical points and invents two facts along the way.
The strengths are real: token-economical behavior, cleanly structured responses, a solid security baseline, serviceable reasoning quality, and formally sound tool integration. The weaknesses are equally real — and more important for deployment than any given average: seven timeouts in 49 tests, critical tail latency, repeated language errors, one hard word-limit violation with an automatic 20% deduction, and above all three hallucinating Tool Use cases. For agentic research or fact-critical workflows, that is too much. For local knowledge work, drafts, initial analyses, and non-critical documentation assistance, Hermes can be a sensible choice — provided a human retains final control.
In short: this model has intelligence, but not iron discipline. Those looking for a free, locally deployable assistant with character will find substance here. Those who need reliability without follow-up review should keep looking.
This evaluation was generated automatically based on the benchmark data. Model used: GPT 4.5 by OpenAI. The raw data and full methodology are documented in the GitHub project.