Tool-use review
· Instruction-Tuned
Deployment Verdict
Conditional deploy: tool execution is strong, but synthesis does not remain stable enough on tool results, and the tool call was not consistently valid throughout the run. The combined finding is good, but for production MCP pipelines it is only reliable with tight response validation.
Tool Execution Profile
Claude Haiku 5.5 demonstrates clear tool intelligence rather than mere routine pattern-matching. On the Web Search & Tool Selection test — which checks whether the model selects web_search over fetch without being prompted — the model reliably identifies the correct access path. This speaks to workable orchestration in dynamic pipelines. On the URL Construction test, which measures autonomous derivation of a target URL followed by a fetch, it performs adequately but not deterministically enough for fragile target systems. P1 scores are strong overall, yet the global finding “Tool-Call valid: false” is relevant: the model is operationally competent, but not protocol-safe enough to be placed behind a production tool infrastructure without review. No retry was required. This points less toward a comprehension problem and more toward isolated format or call weaknesses under load.
Synthesis Fidelity
How well does it condense tool results? Only adequately. Condensation quality is visibly below execution quality. Solid results on HTTP Fetch & Extract and URL Construction & Fetch stand against weak condensations on EU License Research and Multilingual Search & Synthesis. For pipelines that need not only to retrieve but to produce reliable summaries, compliance notes, or management summaries, this is the primary weakness.
Does it stay within tool results or fall back on training? Not alarming, but not confidence-inspiring either. On the honeypot EU License Research test — which checks whether current license restrictions are actually drawn from web sources — the model does not hallucinate. That is the minimum requirement. The weak P2 finding shows, however, that it does not translate the researched content precisely enough into a reliable answer. For compliance-adjacent paths, that is insufficient.
Error Resilience
Well suited for production. On the 404 test, which measures transparent handling of a failed tool call, the model communicates the error openly and does not fabricate page content. Exactly this behavior keeps a pipeline trustworthy. A tool failure thus remains an operational error and does not become a data integrity problem.
Operational Profile
Call 1: 1.09s. MCP latency: 1.25s. Call 2: 5.49s. Total: 46.98s. For a Haiku-class model, the total runtime is long. Cost/run: local. Favorable relative to performance; not exceptionally efficient relative to response fidelity.
Conclusion & Recommendation
Suitable for MCP pipelines with clear tool guidance, a robust post-validation layer, and low tolerance for fabricated content on error cases. Well suited for retrieval, search selection, pre-structuring, and transparent error handling. Not the first choice for compliance summaries, multilingual research condensation, or any pipeline where the final verbal response must carry the same level of trust as the tool call itself.
This analysis was generated automatically based on benchmark data. Model used: GPT-5.4 by OpenAI. The raw data and full methodology are documented in the GitHub project.