Ornith 1.5 35B-A3B

Ornith 1.5 35B-A3B has been the mid-tier model in DeepReinforce’s open Ornith family since August 19, 2026. The MoE activates only around 3B of 35B parameters per token, yet according to the manufacturer it significantly outperforms the similarly sized Qwen 3.6-35B on all coding and agentic benchmarks. Trained with a closed self-improvement loop that jointly optimizes its own tasks, scaffolds, and solutions. License: MIT, fully open and commercially usable.

DeepReinforce Version 1.5 Commercial use permitted MoE 35 B (3 B active) 262 K Context 05/2026 locally tested

  • Open Weights
  • Workstation
  • vLLM
  • Text
  • Vision
  • Batch

Sovereign Risk: LOW DeepReinforce is a US-based research team. The model is released under the permissive MIT license with fully open weights, enabling independent auditing and fully local operation without cloud dependency. Local deployment involves no additional data transmission to the developer.

LLM Model Review

Created on

With an overall score of 75.65%, Ornith 1.5 35B-A3B is no smoke-and-mirrors act but a serious working model with a clear technical signature. Its speed profile is Batch DevOps Expert, and it behaves exactly as that label suggests: not as a frantic chat all-rounder, but as a model built for longer, structured workflows where planning and technical substance matter more than rapid-fire exchanges. The classification fits: primary use case Agentic / Orchestration, size class Workstation, and as a MoE only 3.0 billion active parameters per token against 35 billion total. That explains quite a bit about both its character and its limits. Sovereign Risk: MEDIUM — DeepReinforce is a US-based provider; local deployment means no prompts reach the vendor, but the legal origin remains US-shaped for governance purposes.

Header Scores: Stability and Reliability

Metric Value Rating Analysis
Timeout Rate 7/49 Unreliable The model is unreliable and drops out at a significant rate in practice.
P95 Response Time 154.0 s Critical Extreme tail latency. The model exhibits massive variance and is unsuitable for time-critical processes.

Ornith runs in this report explicitly in Thinking Mode. That matters, because expectations here differ from a terse instruct run: longer reasoning phases, more elaborate derivations, and a certain tendency toward over-explanation are not just tolerated but part of the design. At the same time, this model is listed in its metadata as Coder, Agentic, and MoE. It should not be read as a pure writing assistant but as a technical planner that thinks across code, process, and structure simultaneously. Text-only benchmarks capture only a slice of its multimodal makeup; vision capability is present in the metadata but barely exercised in the current test suite.

Architecture and Character: High Ambition on Limited Active Capacity

The most interesting number about Ornith 1.5 35B-A3B is not 35 but 3. The model formally belongs to the Workstation class, yet as a Mixture-of-Experts architecture it activates only around 3 billion parameters per token. That active capacity is the honest benchmark. What you get is not the sustained force of a densely packed 35B model but a more specialized system that draws its strength from routing, task matching, and efficiency.

That is also why the results feel so characteristic. In modules that live on technical structure, security reasoning, and multi-step problem analysis, Ornith performs strongly. In areas where multiple soft constraints must be satisfied cleanly at the same time — language, length, format, and tone — it becomes noticeably more fragile. That is not a full exoneration. But it is a clean classification: this model prefers to think too much rather than too little, and in doing so it occasionally loses the last mile of instruction discipline.

Speed and Runtime Behavior

Running as a local model on ASUS GX10 / NVIDIA DGX Spark (GB10 Grace Blackwell Superchip, ~115 GB Unified Memory — no practical memory limit for tested model sizes), Ornith shows the typical profile of a thinking-oriented Agentic model: not a sprinter, but not a throughput failure either. The Batch DevOps Expert badge is apt. It signals a model better suited to planned technical workflows than to latency-sensitive dialogues or tight tool loops with hard response-time requirements.

In practice that means: anyone integrating Ornith into an agent stack should think asynchronously. Background jobs, audit runs, security analyses, larger code reviews, and structured problem-solving are good fits. Interactive surfaces where users expect rapid, concise responses in quick succession are not. A second issue compounds this: token economy. In the CLI module Ornith stays within reasonable bounds, and likewise in UX writing. But in the code quality domain it produces an average of 7,359 output tokens against a fleet median of 2,947. That is a factor of 2.5. For a local model this is primarily a latency signal. Ornith solves tasks not only through competence but also through sheer volume.

Code Quality: Technically Strong, but Prone to Long Wind-Ups

In the code quality audit Ornith reaches 79.88%. That is no lucky strike but an expression of genuine technical maturity. In security contexts in particular the model works precisely: SQL injection in multiple variants, type juggling through loose comparisons, path traversal, IDOR, XSS, session fixation, weak reset tokens, cookie-based admin authorization, and implicit attack surfaces such as mail header injection — all identified reliably. More importantly, the proposed fixes are generally not merely abstract but practically actionable.

The qualitative log shows clearly where Ornith convinces. On a PHP security analysis it delivers a clean Markdown table, maintains the German language requirement, keeps table cells concise, and then usefully adds a focus on the requested implicit vulnerabilities. The judge notes 17 vulnerabilities found against 19 in the reference material. Not perfect, but the critical points land. Anyone who works with security reviews in practice knows: the last two minor findings are annoying. What matters is whether a model catches the fatal flaws. Ornith does.

Security is one of the clear strengths of this run overall. The model does not merely name vulnerabilities — it understands their attack paths. It reads like an experienced auditor, not like a glossary on autocomplete. That is worth a great deal, especially for a MoE with only 3 billion active parameters.

The catch is form. In the code domain Ornith tends toward extended internal reasoning chains and longer visible outputs, even when the task could be resolved with far less text. That does not make the answers worse, but it makes them heavier. For developers who want a quick, concise patch suggestion this can be frustrating. For audit or review scenarios it is acceptable.

Reasoning and Logic: Correct, Sound, Not Maximally Didactic

In the logical reasoning module Ornith lands at 75.64%. That is a solid result and consistent with the architecture labels Reasoning and Thinking. Importantly, this run does not show a model that thinks in spectacularly original ways — it shows one that thinks cleanly and reliably. On the classic two-guards puzzle it correctly arrives at the well-known inverse question and explains the logic coherently. The judge praises the substantive correctness but notes that the presentation remains more concise and pedagogically less developed than the reference.

That captures the character fairly precisely. Ornith is dependable in its reasoning but not always elegant. It builds sound inference chains without casting them in the best possible didactic form. Anyone looking for a model that brilliantly unpacks complex logic for non-specialists will find a matter-of-fact engineer here rather than a good teacher.

Precisely because it was tested in Thinking Mode, a degree of verbosity is to be expected. It delivers that. But the added value of that length is not always proportional. The log shows deep internal reasoning while the visible answer then remains comparatively sober. That is not a flaw. It simply shows that Ornith invests more computation than it ultimately returns to the user as visible structure.

Content Transformation and Tool Proximity: Strongly Built, Then Stopped by the Guard Rails

With 79.96% in content transformation, Ornith initially shows a surprisingly strong side. The model can structure production material and enrich it with timestamps, screen annotations, retention hooks, CTAs, and even Easter eggs. That is technically impressive and speaks to its agentic design: it thinks in process, not just in sentences.

Yet this is also where one of the biggest flaws of this run sits. In a video script task within the content domain, the model exceeded the explicit word limit of 900 words by 58%. The system applied an automatic deduction of 16.72 points, or 20% of the achieved partial score. The substantive quality of the answer is therefore irrelevant. The penalty applies regardless.

A language error compounds the same task: Ornith ignored the explicit language instruction and responded in the spoken section in English, despite German being required. The judge describes the answer as structurally strong and production-ready, but linguistically off-target. That is more than a cosmetic flaw. Anyone working with a fixed target language — in editorial, marketing, or customer communications — cannot afford that kind of slip.

And here it becomes structural. The language failure is not an isolated outlier. Across multiple tasks in the content and tool domain the model shows a consistent pattern: when simultaneous constraints on language, length, and format are present, it drops the language requirement first. Beyond the content task already mentioned, a further language error occurred in the tool use domain. For a model with agentic ambitions this is uncomfortable, because agents in particular depend on clear output contracts. A good planner who delivers in the wrong language is still the wrong choice.

UX Writing and Cultural Intelligence: Not Its Natural Habitat

In UX writing Ornith drops to 70.37%. Not a collapse, but a clear deceleration. In one task that required only the rewritten German text, the model delivered a usable, inclusive version — then appended five lines of justification. That was explicitly prohibited. The judge calls the rewrite itself solid but the structural disregard of “output only” severe.

This is typical of technically oriented Thinking models: they solve the problem and then want to explain why they solved it. In a security analysis that is helpful. In microcopy it is an own goal. Ornith does not write badly here. It simply writes too much from its own internal operating manual.

Cultural Intelligence also remains decent but not outstanding. On the positive side: solid command of German and the ability to redirect toxic or exclusionary phrasing into more professional territory. On the negative side: a certain lack of lightness. Where good UX microcopy should feel effortless, Ornith sounds more like a thoroughly revised protocol. It hits the purpose, but rarely the fine tone.

Tool Execution and Agentic Fit: Planning Yes, Format Discipline With Reservations

The tool execution score of 71.67% is solid, but not spotless for a model with an explicitly agentic focus. The result should be read with nuance. Agentic models can be judged more leniently when they do not produce every exact one-liner like a shell specialist, because real agent systems often delegate subtasks. What matters then is planning, decomposition, and strategic structure.

That is exactly where Ornith has substance. It structures tasks sensibly in most cases and operates in multi-step patterns that fit orchestrated pipelines. At the same time, the documented language error in tooluse006 shows that it lacks the final layer of strict instruction compliance. That is not a total failure, but it is a warning signal. Agentic models live by contracts: which tool, which format, which language, which output. When one of those conditions quietly slips out of view, planning quickly turns into rework.

Hallucinations and Content Reliability

This run shows no clear hallucination pattern. The weaknesses are visibly more in format and compliance than in freely invented technical nonsense. Particularly in code, security, and reasoning tasks, Ornith comes across as fact-oriented and remarkably sober. It is less inclined to fill gaps with invention than some more verbally exuberant competitors. That is not a glamorous advantage, but it is a valuable one.

Data Privacy and Data Sovereignty

DeepReinforce does not operate its own cloud service. Ornith is provided as an Open Weights model under the MIT license, so the hosting and data privacy reality rests with the operator or a chosen third-party provider. For European organizations that is good news, since local deployment puts data sovereignty effectively in their own hands.

The origin is not entirely without consequence, however. The calculated Sovereign Risk is MEDIUM. The reason is the US jurisdiction of the vendor. The CLOUD Act remains a relevant legal framework even though DeepReinforce does not operate an inference service itself. In practice that means: local operation does not automatically transfer user data to the vendor, but governance teams should document the US origin of the weights. A GDPR DPA does not arise in the strict sense here, since no provider service is being used. Anyone inferring Ornith through a third-party provider immediately trades that advantage for that provider’s contractual and storage regime.

Conclusion

Ornith 1.5 35B-A3B is a local model with a strong character, genuine technical class, and noticeable opinions about how it works. It excels where security reasoning, code comprehension, structured analysis, and multi-step planning are required. It falters where rigid instructions on language, length, and format must be satisfied simultaneously. That is not a contradiction — it is its profile. Ornith thinks like an ambitious technical contributor who would rather over-justify than under-explain. In some teams that is worth its weight in gold. In others it is simply too slow and too headstrong.

For security reviews, code audits, technical long-form analyses, and asynchronous agent workflows the model is a clear recommendation. For UX microcopy, strictly localized content production, and tightly timed tool execution — only with guard rails, validation, and retries. Across all tests no notable hallucinations — the model prefers to invent little and fails on form rather than facts. On weights provenance there is little cause for concern: the risk is rated low, and the MIT license permits fully local operation without vendor contact. Compared to Ornith 1.0, the 1.5 version feels more modern overall and somewhat sharper in content and code domains, while losing none of the family’s fundamental batch nature. Anyone looking for a quick charmer is in the wrong place. Anyone looking for a thorough technical workhorse should take a close look.

This evaluation was generated automatically based on the benchmark data. Model used: GPT 4.5 by OpenAI. The raw data and full methodology are documented in the GitHub project.