Claude Opus 5

Claude Opus 5 is Anthropic’s Opus flagship as of July 24, 2026, positioned as an everyday model between Opus 4.8 and the more expensive Fable 5. The cloud-only model under US jurisdiction offers 1 million tokens of context, 128,000 tokens of output, and adaptive reasoning control with five effort levels (low/medium/high/xhigh/max). Mid-conversation tool switching without cache loss and a Fast Mode with 2.5× speed round out the offering.

Anthropic Version 5 Commercial use permitted Dense 1000 K Context 05/2026 $5 / $25 per 1M

  • Proprietary
  • Frontier
  • Anthropic
  • Text
  • Vision
  • Agentic Orchestrator
  • Long Context
  • Interactive

Sovereign Risk: MEDIUM Anthropic is a US-based company and subject to the CLOUD Act. The model weights are proprietary and not distributed; no additional risk from weight distribution. Data handling is governed by Anthropic Commercial Terms.

Tool-use profile: 6 assets in detail

Comparison of asset performance (P1/P2/Combined) against the fleet average

Asset performance (radar)

Score Breakdown vs. Fleet Average


Tool-use details

Asset performance, reliability, and runtime profile

CrucibleMark evaluates tool use across 6 independent tests. Click a test name for details.

Reliability

  • Tool Call Valid: No
  • Retry: Not required
  • Hallucination: Not detected

Reliability measures how consistently a model actually executes tool calls: Tool Call Valid schema and format accepted, Retry Required successful only after retry, Hallucination Flag fabricated tools or parameters detected. All three green means production-ready.

Runtime profile

Call 1
2.39
First Request
MCP
1.37
Protocol Latency
Synthesis
16.4
Response Generation
Total
120.95
Sum of All Phases
Token
10412
Input + Output
Cost
$0
Cost per Run

The runtime profile shows the latency and cost metrics for the model run: Call 1 First Request, MCP Protocol Latency, Synthesis Response Generation, Total sum of all phases, supplemented by Token input and output and Cost cost per run.

Tool-use review

Updated · Agentic Orchestrator · Long Context

Deployment Verdict

Conditional deploy: overall performance is solid and no hallucination was detected, but Tool Calls were not consistently valid. For production MCP pipelines this is manageable, provided a strict call validator and fallback paths are in place upstream.

Tool Execution Profile

Claude Opus 5 demonstrates genuine tool intelligence rather than rigid schema-following behavior. On the Web Search & Tool Selection test — which checks whether the model selects web_search over fetch without an explicit hint — the model reliably identifies the correct mode. This speaks to usable orchestration capability in dynamic pipelines. On the URL Construction test, which measures correct URL derivation followed by a fetch, performance is merely solid. The model can often derive the target URL adequately, but not with the precision required for deterministic flows with tight error tolerance.

The critical issue is not tool selection but protocol compliance during execution. Tool Call valid: false means this model should not be handed tool infrastructure without oversight. On the positive side, no retry was required. This reads more like a robustness deficit in call formatting or, in isolated cases, parameterization — not a fundamental comprehension problem.

Synthesis Fidelity

How well does it condense tool results? Well, but not consistently sharp. Condensation remains usable and transparent across most tasks, with clear strength on Tool Failure Handling (404) and solid performance on HTTP Fetch & Extract. The visible weak point is Multilingual Search & Synthesis: cross-language research succeeds, but the German-language consolidation loses accuracy and prioritization. This warrants attention for multilingual compliance, policy, or market-monitoring pipelines.

Does it stay within the tool result or fall back on training? On the Honeypot EU License Research test — which checks whether current license restrictions are answered from web sources rather than training data — the model stays on the safe side. No hallucination detected. This is the central trust signal of this run.

Error Resilience

When tool calls fail, Claude Opus 5 responds in a production-appropriate manner. On the 404 test, which measures transparent error communication rather than fabricated page content, the model does not hallucinate substitute content and marks the failure cleanly. This is acceptable for production tool pipelines.

Operational Profile

Total: 120.95s. Time to first token: 2.39s, MCP latency: 1.37s, second call: 16.40s. Slow for interactive workflows, acceptable for deep agentic runs. Pricing: $5.0/1M input, $25.0/1M output. Not cheap for Frontier-tier, but justifiable where the pipeline benefits from long context and orchestration.

Summary & Recommendation

Suitable for agentic research, routing, and long-context pipelines with a validation layer — particularly where tool selection matters more than millimeter-precise URL construction. Not the first choice for strictly deterministic tool chains, multilingual synthesis with high precision requirements, or low-latency user flows. Teams that implement call validation, schema checks, and clear fallbacks can put it to productive use.

This analysis was generated automatically based on benchmark data. Model used: GPT-5.4 by OpenAI. The raw data and full methodology are documented in the GitHub project.