Hermes 4 14B

Hermes 4 14B as a Q4 quantization of the NousResearch distribution based on Qwen-3, optimized for local assistance and agentic tasks. With 14 billion parameters and a 128,000-token context window, the model runs on resource-constrained hardware and supports hybrid reasoning modes. Fully commercially usable under the Apache 2.0 license.

NousResearch Version 4.0 Commercial use permitted Dense 14 B (14 B active) 128 K Context 09/2024 locally tested

  • Open Weights
  • Desktop
  • llama.cpp
  • Text
  • Community-Quantisierung
  • Instruction-Tuned
  • Interactive

Sovereign Risk: MEDIUM NousResearch is a US-based company; the CLOUD Act is only relevant when using the API, not when running the Open Weights variant locally.

Key metrics

Score · Latency · Cost · Quality

Total Score Silver
67.32
Routine
41.21
Reasoning
26.11

Rank #80

LLM Judge Avg
3.27
100 Coverage
Avg Task Duration
44.76
Interactive
Token Rate
30.3
Output Rate
P95 Latency
139.27
Top 5 %
Total Tokens
70700
Output Volume
Cost per 1K
$0
USD / 1K Requests
Benchmark Cost
$0
Total · 70700 tok

Benchmark modules

10 modules · weighted · vs. model median & top performer

Hermes 4 14B Best model Ø All models
Code Quality 62.5
CLI Benchmark 85.56
Logical Reasoning 64.02
UX Writing 56.65
Documentation 64.83
Content Transform. 73.3
Cultural Intelligence 73.6
Synthesis Quality 45
Tool Execution 90
ToolUse Score 67.21
Benchmark Cost $0

Token efficiency & latency

Consumption per module vs. model median

Token consumption per module

Performance profile