Ornith 1.0 35B

What FP8 block quantization delivers with Ornith-1.0-35B-FP8: an Open Weights MoE with only around 3 of 35 billion active parameters per token runs on a single GPU and brings 262,144 tokens of context, native thinking, and tool calling. DeepReinforce trained the model to learn its own agentic approach rather than working with a fixed rule set. MIT license, commercial use, and fine-tuning without restrictions.

DeepReinforce Version 1.0 Commercial use permitted MoE 35 B (3 B active) 262 K Context 05/2026 locally tested

  • Open Weights
  • Workstation
  • vLLM
  • Text
  • Batch

Sovereign Risk: LOW DeepReinforce is a US-based RL research organization. The model is available on Hugging Face under the MIT license without regional restrictions (79,608 downloads/month). Lineage: Qwen3.5-35B-A3B (hybrid MoE base, Alibaba Cloud) + Gemma 4 → DeepReinforce Ornith-1.0-35B (RL post-training) → official FP8 block quantization (E4M3) by the same author. No Chinese NSL risk, no US CLOUD Act risk when operated locally, as it is a pure Open Weights model with no cloud API requirement.

LLM Model Review

Updated on

With an overall score of 75.79%, Ornith 1.0 35B does not present itself as an accommodating generalist but as a specialized workhorse with a distinct signature. This fits its classification: primarily conceived for agentic use, positioned in the Workstation class, built as an MoE model — and therefore more sensibly measured against its roughly 3 billion active parameters than its 35 billion total. The Speed Profile Badge reads “Batch DevOps Expert”: not a sprinter for frantic chat windows, but a model for longer, planning-intensive workloads. Sovereign Risk: MEDIUM — DeepReinforce is a US-based provider and therefore subject to US law in principle; however, for the locally hosted Open Weights deployment relevant here, the provider itself processes no user requests.

Header Notes: Stability and Reliability

Metric Value Rating Analysis
Timeout Rate 10/49 Unreliable The model is unreliable and drops out significantly often in practice.
P95 Response Time 162.79 s Critical Extreme tail latency. The model shows massive variance and is unsuitable for time-critical processes.

These header notes are the first reality check. Ornith 1.0 35B can appear impressive when given room to operate. But it can just as easily derail a pipeline’s cadence the moment reliability matters more than intellectual depth. For an agentic model, extended planning is not a flaw. Ten dropouts in 49 tests very much is.

Architecture and Character: Not a Conversationalist, but a Planner with a Toolbox

The pre-assigned categorization captures the essence surprisingly well. As a coder and agentic model, Ornith is not primarily meant to formulate charming prose but to decompose problems, identify risks, and hold the line across multi-step tasks. As a thinking model — and in this specific run explicitly tested with thinking mode enabled — it is permitted to be more elaborate and argumentative than a compact instruct model. And that is exactly what it does.

The MoE architecture — “Mixture of Experts” — is important here: only a subset of weights is active per token. This explains part of its character. Ornith does not feel like a raw 35B powerhouse but like a focused specialist with rotating areas of competence. The upside is efficiency of active model capacity. The downside: results are not always uniform. Especially in benchmarks that demand precision, format discipline, and consistent pacing in equal measure, this uneven distribution shows itself without mercy.

Reasoning: Substantively Strong, Not Always Didactically Elegant

In the reasoning module, Ornith 1.0 35B delivers what one expects from a genuine thinking run: not mere short answers, but traceable solution paths. On the classic guard puzzle, the core logic is correct, the conclusion clean, the case distinction sound. The model explains the double-inversion mechanism correctly and remains linguistically clear. This is not a revelation. But it is the kind of clean intellectual discipline that is needed more often in practice than spectacular flashes of insight.

The issue lies not in the correctness of the result but in its communication. Compared to an ideal model solution, the second and third layers of explanation are missing: more visual compression, more alternative formulations, more pedagogical redundancy in the best sense. Ornith solves the task but does not teach it masterfully. That is an important distinction. Those using a model as an analytical partner will be able to live with this. Those looking to generate training or explanatory content directly from it will get correct substance — but not always the most elegant dramaturgy.

No systematic metacognitive failure is evident here. On the contrary: in the present thinking run, Ornith works with the required thinking tags, keeping format and content aligned. For a model in this category, that is not an optional extra but a baseline requirement. At least it meets it.

Code Quality and Security: Where Ornith Shows Its Real Talent

The most interesting part of this model clearly lies in technical auditing. In the code quality task involving security analysis of a PHP script, Ornith operates with genuine expertise: SQL injection, session fixation, path traversal, CSRF, weak tokenization, type juggling, header injection, information leakage. It lands. Above all, the prioritization lands. Ornith does not merely identify individual defects but assesses their severity sensibly and structures the table in a way that is genuinely useful for real review work.

This is where the coder DNA shows itself. The model does not write ornately over-clever prose but attempts to map attack surfaces into a reliable framework. It names direct vulnerabilities as well as implicit risks — such as mail header injection or problematic cookie handling. It does not hallucinate wildly in the process. The errors in judgment lie more in weighting and framing.

Because it is not perfect. The Judge logs show precisely where Ornith leaves points on the table: it does not isolate individual SQL injection vectors sharply enough, particularly in delete and password-reset paths. It conflates direct vulnerability and downstream consequence where a real security review would need to separate them cleanly. Additionally, the synthesis of an attack path is missing. The model solution demonstrates how individual holes form a chain. Ornith instead delivers something closer to an inventory list than an exploit narrative. For an auditor, that is usable. For a red team briefing, it is not yet sharp enough.

There is also the most serious practical finding in this module: the quality of the analysis stands in grotesque contrast to its stability. The module is substantively strong but ran practically into a wall on the reliability side. Security reviews are precisely the kind of task where no one wants to play a lottery. A model that drops out on four out of five corresponding runs resembles a brilliant pentester who only shows up to every fifth appointment. Talent does not excuse that.

CLI and Agentic Behavior: Strategically Useful, Not Always Operationally Precise

The overall characterization as “Batch DevOps Expert” fits. Ornith appears to treat tasks not as individual commands but as small mission plans. This agentic profile is broadly positive in the benchmark because it promotes structure, sequencing, and implicit tool orientation. The model conveys that it does not merely want to respond but to proceed.

That is an asset when planning is what is needed. It is less of an asset when exact directness is required. The ToolUse score remains solid but not outstanding. This suggests that Ornith does not always cleanly reconcile strategic intent with operational precision. For shell tasks, DevOps triage, or multi-step correction runs, that is acceptable. For environments where every command must land immediately, it requires oversight, retries, and — when in doubt — a tighter prompt frame.

Content Transformation: Good Substance, Poor Length Management

In the Content Transformation domain, Ornith exhibits a recurring trait: it often has something sensible to say but says too much of it. In the video script task, the substantive analysis is quite usable. The model identifies missing hooks, timing markers, B-roll cues, emotional anchors, and CTA structure. The actual script is not bad either — functional, in German, implementable, and in places even engaging.

Then comes the question of discipline. And Ornith answers it unnecessarily poorly.

In a task within the Content Transformation domain, the model exceeded the explicit word limit of 900 words by 21%. The system applied an automatic penalty of 20 percent, or 17.00 points. The substantive quality of the response is therefore irrelevant — the penalty applies regardless. This is not a softly interpretable cosmetic flaw but a clear violation of an explicit production constraint.

This finding is more than a one-off lapse. Token efficiency confirms the pattern. In the Content Transformation module, Ornith produces significantly more text than the median of tested models, even though that is precisely where the task demands precision and pacing. For local use, this means primarily more runtime. For agentic workflows, it means something more fundamental: under competing requirements, the model loses track of the word limit first. It thinks broadly where it should be cutting.

UX Writing and Linguistic Control: Competent, but Prone to Persuading Rather Than Compressing

The UX writing result is solid, but the token finding casts a shadow. Ornith produces massively more text here than necessary. That would be forgivable if it were stylistically superior as a result. The benchmark suggests a different profile instead: the model can write, but not with the economy of a truly skilled product copywriter. It explains, justifies, expands. It does not always write to the point.

This is a classic specialist weakness of coding and agentic models. They want to secure context rather than let friction disappear elegantly. In UX microcopy, that is precisely the wrong instinct. A button label is not an architecture meeting. Anyone deploying Ornith for product copy should not let it run freely but work with strict length and format constraints.

Documentation Quality: Robust Middle Ground with Technical Seriousness

In documentation, Ornith shows a comparatively healthy profile. Responses are structured, explanation-oriented, and technically coherent. The model feels less like an author here and more like a conscientious colleague who documents a handover carefully. That is often more valuable than stylistic flair.

The familiar pattern remains visible here as well: more volume than the median, without that automatically translating into more insight. In documentation, this is less fatal than in UX or content work. Those serving technical teams can live with this kind of breadth. Still, somewhat more editorial rigor in the model’s behavior would be desirable. Not every complete answer is already a good one.

Cultural Intelligence: Surprisingly Confident, but Not Fully Idiomatic

One of the more pleasant findings is the cultural and linguistic adaptation. Ornith remains consistently in German throughout the available logs and generally hits the formal register cleanly. In the application letter and register task, for instance, the verdict is clear: functionally good, culturally appropriate, but not always the finest German equivalent in terms of tone. The model understands what is meant. It just does not always hit the most precise register level.

Notably, this weakness does not stem from gross linguistic uncertainty. It arises more from slightly too technical or too neutral phrasing. For a model with a coder and agentic focus, that is almost a compliment. It shows that Ornith is not linguistically blind — it simply has its natural center of gravity elsewhere.

Speed and Efficiency: Batch Is Not a Metaphor Here

Ornith 1.0 35B was evaluated as a local model natively on ASUS GX10 / NVIDIA DGX Spark (GB10 Grace Blackwell Superchip, ~115 GB Unified Memory — no practical memory limit for tested model sizes). The Speed Profile Badge “Batch DevOps Expert” describes the model accurately: generation is not tuned for immediate response but for longer, work-intensive runs. In plain terms: Ornith is moderate to rather slow — not in the sense of unhurried elegance, but with noticeable tail risk on outliers.

This partly fits the architecture. An agentic thinking model is allowed to take longer if it plans cleanly in return. The problem is not the absence of real-time capability. The problem is the combination of high text output and unstable reliability. Token efficiency shows multiple yellow and red zones: Code Quality, Content Transformation, and especially UX Writing run well above the fleet median. Locally, that means primarily more generated tokens, more wait time, more opportunities for dropouts. Ornith thinks a lot. It also tends to talk too much in the process.

Privacy and Data Sovereignty

A dedicated privacy alert is not fundamentally necessary for this model, since Ornith 1.0 35B is operated locally as an Open Weights model and DeepReinforce does not provide its own cloud inference service. The legal origin remains relevant nonetheless: the developer is based in the US, and the calculated Sovereign Risk is MEDIUM. For the specific local deployment scenario, however, this is substantially mitigated, as the hosting and storage question rests entirely with the user and DeepReinforce itself processes no requests.

Conclusion

Ornith 1.0 35B is a serious Workstation-class model with a clear technical identity. It reaches 75.79% not as a smooth all-rounder but as a local specialist for security analysis, code-adjacent structural work, and reasoning-heavy agentic runs. Its MoE architecture with roughly 3 billion active parameters per token is not a marketing detail but the correct lens for evaluation: measured against that, the intellectual yield is respectable.

But this model comes at a cost, and that cost is practical friction. The instability is too high, the tail latency too critical, the verbosity too pronounced. Ornith is often intelligent but not disciplined enough. It delivers substance in technical modules but loses its footing under hard constraints and produces more material than value in text-adjacent tasks. For local coding and audit workflows with human review in the loop, it is nonetheless an attractive package: open, commercially usable, strong on reasoning, and alert on security questions. For unsupervised agent pipelines or time-critical interaction, it is not a good idea in its current state. Across all tests, no notable hallucinations — the model prefers to rarely invent nonsense rather than embarrass itself with creative overconfidence.

This evaluation was generated automatically based on the benchmark data. Model used: GPT 4.5 by OpenAI. The raw data and full methodology are documented in the GitHub project.