Claude Sonnet 5

Claude Sonnet 5 has been the new default for free and pro users since late June 2026. The successor to Sonnet 4.6 brings Adaptive Thinking as default behavior, speed close to Opus, and processes text and images with a 1,000,000-token context window. Safety is significantly improved, with better resistance to prompt injection and fewer hallucinations.

Anthropic Version 5 Commercial use permitted Dense 1000 K Context 01/2026 $2 / $10 per 1M

  • Proprietary
  • Frontier
  • Anthropic
  • Text
  • Vision
  • Long Context
  • Interactive

Sovereign Risk: MEDIUM Anthropic is a US-based provider; relevant risks relate to cloud processing under US law, as no open weights are available.

Tool-use profile: 6 assets in detail

Comparison of asset performance (P1/P2/Combined) against the fleet average

Asset performance (radar)

Score Breakdown vs. Fleet Average


Tool-use details

Asset performance, reliability, and runtime profile

CrucibleMark evaluates tool use across 6 independent tests. Click a test name for details.

Reliability

  • Tool Call Valid: No
  • Retry: Not required
  • Hallucination: Not detected

Reliability measures how consistently a model actually executes tool calls: Tool Call Valid schema and format accepted, Retry Required successful only after retry, Hallucination Flag fabricated tools or parameters detected. All three green means production-ready.

Runtime profile

Call 1
1.74
First Request
MCP
1.48
Protocol Latency
Synthesis
11.44
Response Generation
Total
88.01
Sum of All Phases
Token
10800
Input + Output
Cost
$0
Cost per Run

The runtime profile shows the latency and cost metrics for the model run: Call 1 First Request, MCP Protocol Latency, Synthesis Response Generation, Total sum of all phases, supplemented by Token input and output and Cost cost per run.

Tool-use review

Created · Long Context

Deployment Verdict

Conditional deploy, because tool use is strong and no hallucination was detected, but tool calls were not consistently valid and synthesis quality is too uneven for production knowledge pipelines.

Tool Execution Profile

Claude Sonnet 5 demonstrates genuine tool intelligence, not just rigid call patterns. In the Web Search & Tool Selection test, which requires choosing between search and direct retrieval without an explicit hint, it selects the appropriate tool reliably. This speaks to usable orchestration in open pipelines. In the URL Construction test, which requires deriving the target URL from its own knowledge and then fetching it, it remains usable but not deterministic enough for flows that expect exact endpoints. The main signal is therefore clear: good choice of tool type, weaker precision in concrete execution. The critical issue remains that tool calls were not consistently valid overall. This is not a total failure, but it is an integration risk for MCP pipelines that require strict protocol compliance.

Synthesis Fidelity

How well does it consolidate tool results? Solid, but not reliably precise enough. The P2 score of 66.67 shows that Sonnet 5 usually combines results meaningfully, but loses noticeable accuracy as soon as multilingual or source-faithful consolidation is required. This is particularly visible in Multilingual Search & Synthesis, where the research works but the consolidation in German abstracts too heavily.

Does it stay within the tool result or fall back on training? Mostly yes, with slight reservations. In the Honeypot EU License Research test, which checks whether current license restrictions are answered from web sources rather than training knowledge, it does not hallucinate. That is the important trust signal. However, P2=60 shows that it does not translate the retrieved content into a reliable answer with maximum precision.

Error Resilience

Acceptable for production. In the 404 test, which checks for transparent behavior when a retrieval fails, Sonnet 5 does not fabricate substitute content. It communicates the error rather than delivering fictitious page data. This behavior is precisely what protects tool pipelines from silent data corruption risk.

Operational Profile

Call 1: 1.74s. MCP latency: 1.48s. Call 2: 11.44s. Total: 88.01s.
Price: $2.0/1M input, $10.0/1M output.
Verdict: fast on individual steps, long on the overall run. Moderately priced for Frontier, but not cost-effective given the uneven synthesis.

Summary & Recommendation

Suitable for agentic MCP pipelines where the model needs to select tools, initiate web research, and surface errors cleanly. Not the first choice for compliance-adjacent, multilingual, or heavily consolidating pipelines where every derived formulation must be source-faithful and reproducible. Deployable as an orchestrator with tight output controls, schema validation, and downstream response verification. Without these guardrails, not suitable as a trusted final authority.

This analysis was generated automatically based on benchmark data. Model used: GPT-5.4 by OpenAI. The raw data and full methodology are documented in the GitHub project.