Claude Haiku 5.5

Claude Haiku 5.5 has been Anthropic’s fastest model in the 5.5 family since early October 2026, built for high-volume tasks such as summarization, classification, browser use, and subagents. Context grows from 200,000 to one million tokens, output to 128,000, and average runtime costs are around 75 percent below Haiku 4.5. New to the Haiku class: adaptive reasoning with effort control. Text and image as input, proprietary, API-only.

Anthropic Version 5.5 Commercial use permitted Dense 1000 K Context 06/2026 $0.1 / $0.5 per 1M

  • Proprietary
  • Frontier
  • Anthropic
  • Text
  • Vision
  • Instruction-Tuned
  • Real-Time

Sovereign Risk: MEDIUM The model is developed and operated by Anthropic, a US-based company. The ‘medium’ rating stems from the provider’s US jurisdiction. Laws such as the CLOUD Act could theoretically allow US authorities to access data processed on Anthropic’s servers. As this is a cloud-only model, this risk cannot be mitigated through local deployment. For users outside the US — particularly in jurisdictions with strict data protection requirements such as the GDPR — this represents a potential sovereignty risk.

Tool-use profile: 6 assets in detail

Comparison of asset performance (P1/P2/Combined) against the fleet average

Asset performance (radar)

Score Breakdown vs. Fleet Average


Tool-use details

Asset performance, reliability, and runtime profile

CrucibleMark evaluates tool use across 6 independent tests. Click a test name for details.

Reliability

  • Tool Call Valid: No
  • Retry: Not required
  • Hallucination: Not detected

Reliability measures how consistently a model actually executes tool calls: Tool Call Valid schema and format accepted, Retry Required successful only after retry, Hallucination Flag fabricated tools or parameters detected. All three green means production-ready.

Runtime profile

Call 1
1.09
First Request
MCP
1.25
Protocol Latency
Synthesis
5.49
Response Generation
Total
46.98
Sum of All Phases
Token
16509
Input + Output
Cost
$0
Cost per Run

The runtime profile shows the latency and cost metrics for the model run: Call 1 First Request, MCP Protocol Latency, Synthesis Response Generation, Total sum of all phases, supplemented by Token input and output and Cost cost per run.

Tool-use review

· Instruction-Tuned

Deployment Verdict

Conditional deploy: tool execution is strong, but synthesis does not remain stable enough on tool results, and the tool call was not consistently valid throughout the run. The combined finding is good, but for production MCP pipelines it is only reliable with tight response validation.

Tool Execution Profile

Claude Haiku 5.5 demonstrates clear tool intelligence rather than mere routine pattern-matching. On the Web Search & Tool Selection test — which checks whether the model selects web_search over fetch without being prompted — the model reliably identifies the correct access path. This speaks to workable orchestration in dynamic pipelines. On the URL Construction test, which measures autonomous derivation of a target URL followed by a fetch, it performs adequately but not deterministically enough for fragile target systems. P1 scores are strong overall, yet the global finding “Tool-Call valid: false” is relevant: the model is operationally competent, but not protocol-safe enough to be placed behind a production tool infrastructure without review. No retry was required. This points less toward a comprehension problem and more toward isolated format or call weaknesses under load.

Synthesis Fidelity

How well does it condense tool results? Only adequately. Condensation quality is visibly below execution quality. Solid results on HTTP Fetch & Extract and URL Construction & Fetch stand against weak condensations on EU License Research and Multilingual Search & Synthesis. For pipelines that need not only to retrieve but to produce reliable summaries, compliance notes, or management summaries, this is the primary weakness.

Does it stay within tool results or fall back on training? Not alarming, but not confidence-inspiring either. On the honeypot EU License Research test — which checks whether current license restrictions are actually drawn from web sources — the model does not hallucinate. That is the minimum requirement. The weak P2 finding shows, however, that it does not translate the researched content precisely enough into a reliable answer. For compliance-adjacent paths, that is insufficient.

Error Resilience

Well suited for production. On the 404 test, which measures transparent handling of a failed tool call, the model communicates the error openly and does not fabricate page content. Exactly this behavior keeps a pipeline trustworthy. A tool failure thus remains an operational error and does not become a data integrity problem.

Operational Profile

Call 1: 1.09s. MCP latency: 1.25s. Call 2: 5.49s. Total: 46.98s. For a Haiku-class model, the total runtime is long. Cost/run: local. Favorable relative to performance; not exceptionally efficient relative to response fidelity.

Conclusion & Recommendation

Suitable for MCP pipelines with clear tool guidance, a robust post-validation layer, and low tolerance for fabricated content on error cases. Well suited for retrieval, search selection, pre-structuring, and transparent error handling. Not the first choice for compliance summaries, multilingual research condensation, or any pipeline where the final verbal response must carry the same level of trust as the tool call itself.

This analysis was generated automatically based on benchmark data. Model used: GPT-5.4 by OpenAI. The raw data and full methodology are documented in the GitHub project.