Hermes 4.3 36B (Thinking)
Hermes 4.3 36B (NVFP4) is the compressed variant of NousResearch’s Open Weights model based on ByteDance Seed-OSS-36B, fully decentralized post-trained via the Psyche network. Fireworks compressed the BF16 checkpoint to NVFP4, significantly reducing the memory footprint of the weights. 36 billion dense parameters, 512,000 tokens of native context, optional thinking mode, and an Apache 2.0 license for free commercial use.
- Open Weights
- Server
- VSPK
- Text
- 36B
- NVFP4
- Compressed-Tensors
- 512K-Context
- Long Context
- Agentic Orchestrator
- Batch
Sovereign Risk: MEDIUM Firworks/Hermes-4.3-36B-nvfp4 is an NVFP4 quantization (LLM-Compressor, long-seq calibration on Rombo-Org/Optimized_Reasoning) of the NousResearch base weights (Hermes-4.3-36B on ByteDance Seed-OSS-36B-Base, decentrally post-trained via Psyche). Apache-2.0 license, publicly available on HuggingFace. NousResearch is subject to the US CLOUD Act; with fully local vLLM deployment there is no cloud data egress, risk remains moderate.
Key metrics
Score · Latency · Cost · Quality
- Total Score Silver
- 68.02
- Routine
- 41.43
- Reasoning
- 26.59
- LLM Judge Avg
- 3.35 / 5
- 100 Coverage
- Avg Task Duration
- 62.15s
- Batch
- Token Rate
- 13.26tok/s
- Output Rate
- P95 Latency
- 160.89s
- Top 5 %
- Total Tokens
- 57400
- Output Volume
- Cost per 1K
- $0
- USD / 1K Requests
- Benchmark Cost
- $0
- Total · 57400 tok
Benchmark modules
10 modules · weighted · vs. model median & top performer
Token efficiency & latency
Consumption per module vs. model median