Hermes 4.3 36B

Hermes 4.3 36B (NVFP4) is the compressed variant of NousResearch’s Open Weights model based on ByteDance Seed-OSS-36B, fully decentralized post-trained via the Psyche network. Fireworks compressed the BF16 checkpoint to NVFP4, significantly reducing the memory footprint of the weights. 36 billion dense parameters, 512,000 tokens of native context, optional thinking mode, and an Apache 2.0 license for free commercial use.

NousResearch Version 4.3 Commercial use permitted Dense 36 B 512 K Context 01/2025 locally tested

  • Open Weights
  • Server
  • VSPK
  • Text
  • 36B
  • NVFP4
  • Compressed-Tensors
  • 512K-Context
  • Long Context
  • Agentic Orchestrator
  • Batch

Sovereign Risk: MEDIUM Firworks/Hermes-4.3-36B-nvfp4 is an NVFP4 quantization (LLM-Compressor, long-seq calibration on Rombo-Org/Optimized_Reasoning) of the NousResearch base weights (Hermes-4.3-36B on ByteDance Seed-OSS-36B-Base, decentrally post-trained via Psyche). Apache-2.0 license, publicly available on HuggingFace. NousResearch is subject to the US CLOUD Act; with fully local vLLM deployment there is no cloud data egress, risk remains moderate.

Key metrics

Score · Latency · Cost · Quality

Total Score Silver
67.14
Routine
41.19
Reasoning
25.95

Rank #86

LLM Judge Avg
3.24
100 Coverage
Avg Task Duration
84.81
Batch
Token Rate
12.71
Output Rate
P95 Latency
191.16
Top 5 %
Total Tokens
56700
Output Volume
Cost per 1K
$0
USD / 1K Requests
Benchmark Cost
$0
Total · 56700 tok

Benchmark modules

10 modules · weighted · vs. model median & top performer

Hermes 4.3 36B Best model Ø All models
Code Quality 66.8
CLI Benchmark 82.22
Logical Reasoning 62.23
UX Writing 60.85
Documentation 62.58
Content Transform. 69.42
Cultural Intelligence 74
Synthesis Quality 43.33
Tool Execution 90
ToolUse Score 66.54
Benchmark Cost $0

Token efficiency & latency

Consumption per module vs. model median

Token consumption per module

Performance profile