Claude Opus 4.7

Since mid-April 2026, Claude Opus 4.7 has been Anthropic’s most capable model, designed for coding, agentic loops, and complex reasoning. The xhigh effort level pushes Extended Thinking to maximum depth; the context window spans one million tokens with no surcharge for long contexts.

Anthropic Version 4.7 Commercial use permitted Dense 1000 K Context 01/2026 $5 / $25 per 1M

  • Proprietary
  • Frontier
  • API
  • Text
  • Vision
  • Long Context
  • Interactive

Sovereign Risk: MEDIUM Anthropic is a US-based company and subject to the CLOUD Act. Closed-source model with first-party safety filters (Anthropic Safety); no weights available.

Tool-use profile: 6 assets in detail

Comparison of asset performance (P1/P2/Combined) against the fleet average

Asset performance (radar)

Score Breakdown vs. Fleet Average


Tool-use details

Asset performance, reliability, and runtime profile

CrucibleMark evaluates tool use across 6 independent tests. Click a test name for details.

Reliability

  • Tool Call Valid: No
  • Retry: Not required
  • Hallucination: Not detected

Reliability measures how consistently a model actually executes tool calls: Tool Call Valid schema and format accepted, Retry Required successful only after retry, Hallucination Flag fabricated tools or parameters detected. All three green means production-ready.

Runtime profile

Call 1
1.9
First Request
MCP
0.97
Protocol Latency
Synthesis
16.69
Response Generation
Total
117.42
Sum of All Phases
Token
16071
Input + Output
Cost
$0
Cost per Run

The runtime profile shows the latency and cost metrics for the model run: Call 1 First Request, MCP Protocol Latency, Synthesis Response Generation, Total sum of all phases, supplemented by Token input and output and Cost cost per run.

Tool-use review

Created · Long Context

Deployment Verdict

Created on: 14.06.2026, 16:20:09

Conditional deploy, because tool execution is reliable and protocol-compliant, but synthesis quality is still too inconsistent for production knowledge pipelines.

Tool Execution Profile

Claude Opus 4.7 is robust at the tool layer. Tool calls were valid, no retry was needed, and it shows no signs of MCP format instability. The most important signal is tool selection: in the Web Search & Tool Selection test, which requires distinguishing between search and direct retrieval without an explicit hint, it selects the correct tool confidently. This argues against rigid pattern behavior and in favor of genuine situational assessment.

Weaker is precision on the URL Construction test, which requires deriving the target URL from its own knowledge and then retrieving it correctly. Performance here is sufficient for workable execution, but not for fully deterministic pipelines. In clear search and retrieval chains the model is strong. In flows where it must reconstruct target addresses itself, guardrails or validation steps should be placed upstream.

Synthesis Fidelity

How well does it condense tool results? Solidly, but not consistently at Frontier level. It performs well on HTTP Fetch & Extract and on Tool Failure Handling (404), where it cleanly summarizes retrieved content. Noticeably weaker is Multilingual Search & Synthesis, where condensation across language boundaries loses measurable precision. This is not an execution failure, but a quality risk for international research or policy pipelines.

Does it stay within the tool result or fall back on training data? Predominantly yes, and that is the more important finding. In the Honeypot EU License Research test, which checks whether current licensing restrictions are answered from web sources rather than training knowledge, it remained verifiably grounded in the retrieved material. The P2 score of 60 indicates that condensation was not clean enough. The decisive point, however: no hallucination, no covert fallback to stale knowledge.

Error Resilience

When tool calls fail, the model is production-ready. In the Tool Failure Handling (404) test, which checks for transparent communication rather than fabricated fallback content, it names the error openly and does not hallucinate page content. Exactly this behavior is acceptable in production pipelines.

Operational Profile

Total 112.66s. Individual calls 2.45s and 15.04s, MCP latency 1.29s. Slow for interactive flows. Cost per run 0.191580 USD. Expensive relative to a merely good rather than very good overall performance.

Summary & Recommendation

Suitable for agentic pipelines with multiple tool steps, high hallucination-safety requirements, and tolerance for latency and cost. Particularly well-suited for research, fetch-driven analysis, and workflows where errors must be caught transparently. Not the first choice for multilingual knowledge pipelines requiring condensation, cost-sensitive high-volume routes, or strictly deterministic flows with self-constructed URL logic and no additional validation.

This analysis was generated automatically based on benchmark data. Model used: GPT 4.5 by OpenAI. The raw data and full methodology are documented in the GitHub project.