Ornith 1.0 35B
What FP8 block quantization delivers with Ornith-1.0-35B-FP8: an Open Weights MoE with only around 3 of 35 billion active parameters per token runs on a single GPU and brings 262,144 tokens of context, native thinking, and tool calling. DeepReinforce trained the model to learn its own agentic approach rather than working with a fixed rule set. MIT license, commercial use, and fine-tuning without restrictions.
- Open Weights
- Workstation
- vLLM
- Text
- Batch
Sovereign Risk: LOW DeepReinforce is a US-based RL research organization. The model is available on Hugging Face under the MIT license without regional restrictions (79,608 downloads/month). Lineage: Qwen3.5-35B-A3B (hybrid MoE base, Alibaba Cloud) + Gemma 4 → DeepReinforce Ornith-1.0-35B (RL post-training) → official FP8 block quantization (E4M3) by the same author. No Chinese NSL risk, no US CLOUD Act risk when operated locally, as it is a pure Open Weights model with no cloud API requirement.
Key metrics
Score · Latency · Cost · Quality
- Total Score Silver
- 75.79
- Routine
- 46.56
- Reasoning
- 29.23
- LLM Judge Avg
- 3.92 / 5
- 100 Coverage
- Avg Task Duration
- 79.56s
- Batch
- Token Rate
- 35.53tok/s
- Output Rate
- P95 Latency
- 161.52s
- Top 5 %
- Total Tokens
- 176500
- Output Volume
- Cost per 1K
- $0
- USD / 1K Requests
- Benchmark Cost
- $0
- Total · 176500 tok
Benchmark modules
10 modules · weighted · vs. model median & top performer
Token efficiency & latency
Consumption per module vs. model median