Ornith 1.0 35B

What FP8 block quantization delivers with Ornith-1.0-35B-FP8: an Open Weights MoE with only around 3 of 35 billion active parameters per token runs on a single GPU and brings 262,144 tokens of context, native thinking, and tool calling. DeepReinforce trained the model to learn its own agentic approach rather than working with a fixed rule set. MIT license, commercial use, and fine-tuning without restrictions.

DeepReinforce Version 1.0 Commercial use permitted MoE 35 B (3 B active) 262 K Context 05/2026 locally tested

  • Open Weights
  • Workstation
  • vLLM
  • Text
  • Batch

Sovereign Risk: LOW DeepReinforce is a US-based RL research organization. The model is available on Hugging Face under the MIT license without regional restrictions (79,608 downloads/month). Lineage: Qwen3.5-35B-A3B (hybrid MoE base, Alibaba Cloud) + Gemma 4 → DeepReinforce Ornith-1.0-35B (RL post-training) → official FP8 block quantization (E4M3) by the same author. No Chinese NSL risk, no US CLOUD Act risk when operated locally, as it is a pure Open Weights model with no cloud API requirement.

Key metrics

Score · Latency · Cost · Quality

Total Score Silver
75.79
Routine
46.56
Reasoning
29.23

Rank #15

LLM Judge Avg
3.92
100 Coverage
Avg Task Duration
79.56
Batch
Token Rate
35.53
Output Rate
P95 Latency
161.52
Top 5 %
Total Tokens
176500
Output Volume
Cost per 1K
$0
USD / 1K Requests
Benchmark Cost
$0
Total · 176500 tok

Benchmark modules

10 modules · weighted · vs. model median & top performer

Ornith 1.0 35B Best model Ø All models
Code Quality 76.92
CLI Benchmark 90.67
Logical Reasoning 73.96
UX Writing 77.15
Documentation 78.65
Content Transform. 71.79
Cultural Intelligence 81.44
Synthesis Quality 62.5
Tool Execution 80.83
ToolUse Score 63.21
Benchmark Cost $0

Token efficiency & latency

Consumption per module vs. model median

Token consumption per module

Performance profile