Tool-use review
Updated
Deployment Verdict
Conditional deploy, because tool execution is often usable, but the model does not work reliably enough in tool selection and result synthesis for autonomous MCP pipelines. The absence of hallucination findings mitigates the risk; the invalid tool calls and the only moderate overall impression do not.
Tool Execution Profile
Ornith 1.0 35B can execute tools, but not with consistent protocol fidelity. P1 of 80.83 indicates basic operational capability. The problem lies not in simple retrieval but in selecting the right tool. On the Web Search & Tool Selection test — which checks whether the model chooses search over fetch without being prompted — the model clearly falls short with P1 35. This argues against robust tool intelligence and suggests a fixed pattern instead: known URLs or direct fetch paths work; open research paths perform noticeably worse. The URL Construction & Fetch test, which evaluates deriving a target URL from internal knowledge, confirms this with P1 80. HTTP Fetch & Extract and Multilingual Search & Synthesis also show that the model processes available sources adequately when the access path is already clear. For dynamic tool routers, that is too weak. For pre-structured pipelines with a narrow tool selection, it is usable.
Synthesis Fidelity
How well does it consolidate tool results? Only to a limited degree. P2 of 62.50 is the actual bottleneck. Ornith extracts relevant facts often enough, but does not consistently condense them into precise, decision-ready answers. This is most visible in EU License Research with P2 40 — even though the tool access itself succeeds — and in Tool Failure Handling (404), also with P2 40.
Does it stay within the tool result or fall back on training data? Here the trust verdict is better than the synthesis quality. In the honeypot EU License Research — which checks whether current license restrictions are actually answered from web sources rather than training knowledge — no hallucination was detected. That is a strong signal for compliance-adjacent pipelines. It does not mean the answer quality is high. It only means the model does not undermine the infrastructure with fabricated source content.
Error Resilience
On the 404 test, which measures transparent handling of failed tool calls, Ornith does not hallucinate substitute content. That is a productively relevant positive. However, the communication of the error remains too weakly synthesized and is not always decision-oriented. Acceptable for supervised systems. Too unreliable for fully autonomous agent loops.
Sovereignty Profile
Locally deployable, MIT-licensed, and usable without cloud dependency. At the same time, it sits -1.85 points below the fleet average of 67.58. That is close enough to the fleet average to be seriously relevant for sovereign environments, but not strong enough to offset quality deficits through operational freedom.
Conclusion & Recommendation
Suitable for local, sovereign MCP pipelines with fixed tool paths, human sign-off, and clear fallbacks. Not suitable as an autonomous orchestrator that must reliably choose between search, fetch, and error handling on its own. If you are looking for a local model for document-adjacent research, multilingual synthesis, and controlled tool use, Ornith can deliver. If the pipeline demands independent tool selection and robust final synthesis, you should deploy it only behind guardrails and with a supervisor in place.
This analysis was generated automatically based on benchmark data. Model used: GPT 4.5 by OpenAI. The raw data and full methodology are documented in the GitHub project.