Hermes 4 14B
Hermes 4 14B as a Q4 quantization of the NousResearch distribution based on Qwen-3, optimized for local assistance and agentic tasks. With 14 billion parameters and a 128,000-token context window, the model runs on resource-constrained hardware and supports hybrid reasoning modes. Fully commercially usable under the Apache 2.0 license.
- Open Weights
- Desktop
- llama.cpp
- Text
- Community-Quantisierung
- Instruction-Tuned
- Interactive
Sovereign Risk: MEDIUM NousResearch is a US-based company; the CLOUD Act is only relevant when using the API, not when running the Open Weights variant locally.
Key metrics
Score · Latency · Cost · Quality
- Total Score Silver
- 67.32
- Routine
- 41.21
- Reasoning
- 26.11
- LLM Judge Avg
- 3.27 / 5
- 100 Coverage
- Avg Task Duration
- 44.76s
- Interactive
- Token Rate
- 30.3tok/s
- Output Rate
- P95 Latency
- 139.27s
- Top 5 %
- Total Tokens
- 70700
- Output Volume
- Cost per 1K
- $0
- USD / 1K Requests
- Benchmark Cost
- $0
- Total · 70700 tok
Benchmark modules
10 modules · weighted · vs. model median & top performer
Token efficiency & latency
Consumption per module vs. model median