Claude Sonnet 4.6

Where the Opus class is too expensive, Claude Sonnet 4.6 steps in: coding, computer use, and agentic workflows at near-Opus level, at the lower Sonnet price. The model operates with adaptive thinking in three effort levels, processes text, images, and PDF documents, and offers a context window of one million tokens, generally available since March 2026.

Anthropic Version 4.6 Commercial use permitted Dense 1000 K Context 08/2025 $3 / $15 per 1M

  • Proprietary
  • Frontier
  • Anthropic
  • Text
  • Vision
  • Long Context
  • Interactive

Sovereign Risk: MEDIUM Anthropic is a US-based provider; relevant risks relate to cloud processing under US law, as no open weights are available.

Tool-use profile: 6 assets in detail

Comparison of asset performance (P1/P2/Combined) against the fleet average

Asset performance (radar)

Score Breakdown vs. Fleet Average


Tool-use details

Asset performance, reliability, and runtime profile

CrucibleMark evaluates tool use across 6 independent tests. Click a test name for details.

Reliability

  • Tool Call Valid: No
  • Retry: Not required
  • Hallucination: Detected

Reliability measures how consistently a model actually executes tool calls: Tool Call Valid schema and format accepted, Retry Required successful only after retry, Hallucination Flag fabricated tools or parameters detected. All three green means production-ready.

Runtime profile

Call 1
52.53
First Request
MCP
1.22
Protocol Latency
Synthesis
17.17
Response Generation
Total
425.48
Sum of All Phases
Token
8521
Input + Output
Cost
$0
Cost per Run

The runtime profile shows the latency and cost metrics for the model run: Call 1 First Request, MCP Protocol Latency, Synthesis Response Generation, Total sum of all phases, supplemented by Token input and output and Cost cost per run.

Tool-use review

Created · Long Context

Deployment Verdict

Do not deploy in untrusted tool pipelines: hallucinations were detected, tool calls were not consistently valid, and the overall impression remains only moderate despite serviceable tool execution. For production MCP infrastructure, this is a trust failure, not merely a quality deduction.

Tool Execution Profile

Tool selection is fundamentally strong. In the Web Search & Tool Selection test — which checks whether the model distinguishes between search and direct retrieval without an explicit hint — the model confidently picks the correct tool. This argues against a rigid pattern and in favor of genuine orchestration logic. HTTP Fetch & Extract is also solid, with high precision.

Operational execution is weaker. In the URL Construction test, which requires deriving the correct target URL from model knowledge and then performing a clean fetch, performance is serviceable but not deterministic enough for sensitive pipelines. Added to this is the tool_call_valid: false signal. This means: planning is often correct, but the handoff to tooling does not remain consistently protocol-clean. The absence of any required retry argues less for a mere formatting issue and more for substantive or execution-specific inconsistency.

Synthesis Fidelity

How well does it consolidate tool results? Inconsistently. It demonstrates good synthesis on HTTP Fetch & Extract and on Multilingual Search & Synthesis. But as soon as a task demands more robust consolidation or a concise compliance response, synthesis quality drops noticeably. The P2 score of 51.67 is not a total failure, but too volatile for workflows in which the model’s response serves as a reliable working basis.

Does it stay within tool output or fall back on training data? No. In the Honeypot EU License Research test — which checks whether current license restrictions are answered from web sources rather than training knowledge — the model hallucinates. This is a security risk. Once a model outputs fabricated or simulated current facts as a tool result, the entire tool infrastructure loses its purpose.

Error Resilience

In the 404 test, which checks for transparent handling of failed tool calls, the model hallucinates page content despite the error. This is production-critical without exception. An acceptable response would be a clear error message indicating the failed retrieval. An orchestrating model must not supply fabricated fallback content.

Operational Profile

Call 1: 52.53s. Call 2: 17.17s. MCP latency: 1.22s. Total: 425.48s. Slow for this level of performance. Cost/run: local. Price per model card: $3.0/1M input, $15.0/1M output. Not cheap enough at Frontier level to offset the reliability risks.

Conclusion & Recommendation

Suitable at most for supervised research and drafting pipelines where a human reviews every tool-assisted claim. Not suitable for compliance, license review, incident analysis, automated web research with result forwarding, or any other MCP workflows in which tool responses are treated as factual ground truth. The orchestration appears intelligent, but the model does not reliably maintain the boundary between tool findings and fabricated content.

This analysis was generated automatically based on benchmark data. Model used: GPT-5.4 by OpenAI. The raw data and full methodology are documented in the GitHub project.