Claude Opus 5.5

Claude Opus 5.5 is Anthropic’s new Frontier model: as of September 22, 2026, it replaces Opus 5 and is claimed by the manufacturer to achieve Fable-5.1-level performance at 40 percent lower typical workload costs. Adaptive Thinking is active by default and controllable via effort level; the context window holds one million tokens with 128,000 output tokens. The reasoning classifier errors from Opus 5 have been fixed; only the metacognition refusal remains.

Anthropic Version 5.5 Commercial use permitted Dense 1000 K Context 06/2026 $5 / $25 per 1M

  • Proprietary
  • Frontier
  • Anthropic
  • Text
  • Vision
  • Agentic Orchestrator
  • Long Context
  • Interactive

Sovereign Risk: TODO TODO

Tool-use profile: 6 assets in detail

Comparison of asset performance (P1/P2/Combined) against the fleet average

Asset performance (radar)

Score Breakdown vs. Fleet Average


Tool-use details

Asset performance, reliability, and runtime profile

CrucibleMark evaluates tool use across 6 independent tests. Click a test name for details.

Reliability

  • Tool Call Valid: No
  • Retry: Not required
  • Hallucination: Detected

Reliability measures how consistently a model actually executes tool calls: Tool Call Valid schema and format accepted, Retry Required successful only after retry, Hallucination Flag fabricated tools or parameters detected. All three green means production-ready.

Runtime profile

Call 1
2.85
First Request
MCP
1.21
Protocol Latency
Synthesis
14.03
Response Generation
Total
108.54
Sum of All Phases
Token
8443
Input + Output
Cost
$0
Cost per Run

The runtime profile shows the latency and cost metrics for the model run: Call 1 First Request, MCP Protocol Latency, Synthesis Response Generation, Total sum of all phases, supplemented by Token input and output and Cost cost per run.

Tool-use review

Created · Agentic Orchestrator · Long Context

Tool-use profile

Claude Opus 5.5 achieves a Combined Score of 76.9 (Good) in the tool-use benchmark: P1 Execution 90, P2 Synthesis 62.5, fleet average 68.2.

Strongest test: Tool Failure Handling (404) (64.3). Weakest test: Web Search & Tool Selection (91). The spread between these two tests is -26.7 points.

Reliability status: Tool Call Valid No, Retry Not required, Hallucination Detected.

This data-driven auto-review is compiled from the available tool-use benchmark data. Once a detailed LLM-generated analysis (GPT-5.4) is available, it will automatically replace this template. The raw data and full methodology are documented in the GitHub project.