Claude Opus 4.6

Anthropic’s frontier model for complex agent tasks: Claude Opus 4.6 processes text and image inputs with a standard context window of 200,000 tokens, expandable to one million tokens for long workflows. The model supports tool calls and Extended Thinking for maximum reasoning depth.

Anthropic Version 4.6 Commercial use permitted Dense 1000 K Context 01/2025 $5 / $25 per 1M

  • Proprietary
  • Frontier
  • API
  • Text
  • Vision
  • Agentic Orchestrator
  • Long Context
  • Interactive

Sovereign Risk: MEDIUM Anthropic is a US-based company and subject to the CLOUD Act; model weights are not publicly accessible.

Tool-use profile: 6 assets in detail

Comparison of asset performance (P1/P2/Combined) against the fleet average

Asset performance (radar)

Score Breakdown vs. Fleet Average


Tool-use details

Asset performance, reliability, and runtime profile

CrucibleMark evaluates tool use across 6 independent tests. Click a test name for details.

Reliability

  • Tool Call Valid: No
  • Retry: Not required
  • Hallucination: Detected

Reliability measures how consistently a model actually executes tool calls: Tool Call Valid schema and format accepted, Retry Required successful only after retry, Hallucination Flag fabricated tools or parameters detected. All three green means production-ready.

Runtime profile

Call 1
17.77
First Request
MCP
0.81
Protocol Latency
Synthesis
22.56
Response Generation
Total
246.81
Sum of All Phases
Token
20782
Input + Output
Cost
$0
Cost per Run

The runtime profile shows the latency and cost metrics for the model run: Call 1 First Request, MCP Protocol Latency, Synthesis Response Generation, Total sum of all phases, supplemented by Token input and output and Cost cost per run.

Tool-use review

· Agentic Orchestrator · Long Context

Deployment Verdict

Conditional deploy, because tool execution is strong and protocol-compliant, but the detected hallucination in the honeypot limits confidence in fact-critical MCP operation.

Tool Execution Profile

Claude Opus 4.6 operates clearly above production level on the execution side. Tool calls were valid, retry was not required, and the P1 score demonstrates robust MCP behavior. The tool selection is the decisive factor: in the Web Search & Tool Selection test, which requires the model to distinguish between search and direct fetch without an explicit hint, the model consistently chose the correct tool. This points to genuine orchestration intelligence rather than rigid fetch-first behavior. In the URL Construction test, which requires correct target URLs from the model’s own knowledge, performance was serviceable but not fully deterministic. The pattern is clear: strong decision-making about which tool is needed, somewhat less precision in autonomous target addressing. For dynamic tool pipelines, this is a good profile.

Synthesis Fidelity

How well does it condense tool results? Solid, but not consistent enough for high-trust workloads. Strong condensation on HTTP Fetch & Extract and on Multilingual Search & Synthesis shows that the model can cleanly aggregate structured web content. However, the overall P2 score is pulled down by notable outliers. Particularly in EU License Research and in Web Search & Tool Selection, the substantive synthesis fell off noticeably.

Does it stay within tool results or fall back on training data? This is the central risk. In the honeypot EU License Research, which tests whether current license restrictions are genuinely drawn from web sources, the model fell back on unverified content. This is not merely a quality deficiency — it is a security risk. When a model outputs fabricated or prior-knowledge-based statements as the result of a tool chain, it undermines the reliability of the entire infrastructure.

Error Resilience

The model behaves in a production-appropriate manner when tools fail. In the 404 test, which checks for transparent error handling rather than fabricated fallback content, it communicated the failure cleanly and did not hallucinate page content. This is an important positive finding for real-world MCP pipelines.

Operational Profile

14.39s and 16.63s on the main calls, 1.17s MCP latency, 193.11s per run total. That is slow.
0.273305 USD per run. That is expensive.
Justifiable given the execution strength. Demanding given the synthesis fidelity on fact-critical paths.

Conclusion & Recommendation

Suitable for agentic pipelines with strong tool orchestration, multi-step web research, multilingual processing, and transparent error handling. Not suitable as an uncontrolled terminal instance in compliance, policy, licensing, or other fact-critical flows where tool results must remain strictly sourced. Deploy only with hard guardrails: source binding, constrain responses to tool-backed evidence, and downstream verification for every normative or current-state claim.

This analysis was generated automatically based on benchmark data. Model used: GPT 4.5 by OpenAI. The raw data and full methodology are documented in the GitHub project.