Claude Sonnet 4.6

Where the Opus class is too expensive, Claude Sonnet 4.6 steps in: coding, computer use, and agentic workflows at near-Opus level, at the lower Sonnet price. The model operates with adaptive thinking in three effort levels, processes text, images, and PDF documents, and offers a context window of one million tokens, generally available since March 2026.

Anthropic Version 4.6 Commercial use permitted Dense 1000 K Context 08/2025 $3 / $15 per 1M

  • Proprietary
  • Frontier
  • Anthropic
  • Text
  • Vision
  • Long Context
  • Interactive

Sovereign Risk: MEDIUM Anthropic is a US-based provider; relevant risks relate to cloud processing under US law, as no open weights are available.

LLM Model Review

Created on · Long Context

Claude Sonnet 4.6 achieves an overall score of 79.82% and carries the speed profile badge Interactive DevOps Expert in the Leaderboard. That fits the character of this model remarkably well: a Frontier generalist from the Anthropic API, densely built, clearly optimized for agentic workflows, with multimodal capability and a very large context window, tested in standard mode because there is no Thinking toggle for this cloud model in the benchmark. It does not come across as a sprinter at any cost, but as a system that weights structure, planning, and professional seriousness above show. Sovereign Risk: HIGH — Anthropic is a US provider, subject to the CLOUD Act, and according to the Vendor Card, processing takes place in the United States.

Header Notes: Stability and Reliability

Metric Value Rating Analysis
Timeout Rate 0/49 Stable The model ran absolutely stable and reliably throughout testing.
P95 Response Time 86.21 s Problematic Significant outliers that interrupt workflow.

In practice, this means: the Anthropic API did not fail during testing, but it scatters. Anyone deploying Claude Sonnet 4.6 in interactive agent chains gets reliability, but must live with noticeable wait pockets. For a model with Thinking architecture and optionally deepened reasoning, that is not entirely surprising — but it remains a product characteristic, not a law of nature.

Architecture and Classification

The pre-assigned category here is not merely a label but an explanation. Claude Sonnet 4.6 simultaneously carries the markers Thinking, Vision-Capable, Agentic, Long-Context, and Thinking-Optional. In the actual test run, the mode was set to n/a — meaning standard behavior of a commercial cloud model without a separate Thinking toggle in the benchmark. Even so, it is evident that more is happening internally here than with a straightforward instruct model. Responses are often methodically structured, rarely hurried, and almost always carried by a quiet planning logic.

Also important is the curated classification: primary use case is agentic, the size class is Frontier, the architecture is dense. That sets the bar high. A Frontier model from the Anthropic API must not merely shine occasionally — it must deliver across a broad front. At the same time, this architecture is designed for tool use, task decomposition, and longer contexts. That explains why Claude Sonnet 4.6 performs very strongly in planning, CLI, and structured analysis, while occasionally taking one step too many with strict format requirements. It is not a model that executes blindly. It wants to understand the task first. Often that is a virtue. Sometimes it is the moment where precision tips into verbosity.

The fact that it is vision-capable and can process images and PDFs according to the Model Card remains only partially visible in this text-heavy benchmark. The text performance therefore shows only a cross-section of the actual product character. Anyone looking for multimodal document work should not read this model solely by its text scores. Anyone demanding purely textual precision, on the other hand, gets a fairly honest position fix here.

Performance and Cost Profile

The badge Interactive DevOps Expert signals not raw top speed but a practice-oriented profile for technical dialogues, shell-adjacent workflows, and multi-step assistance in semi-interactive pipelines. Qualitatively, the generation speed places it in the upper interactive range. Claude Sonnet 4.6 does not feel like a batch writer, but neither like a nervous real-time model that cuts every corner to finish fast.

Pricing sits at $3.0 per million input tokens and $15.0 per million output tokens. For a proprietary Frontier model, that is not an outrageous price, but not an impulse purchase either. That is precisely why token discipline matters. And here the picture is mixed: overall the model stays within all budgets, but not elegantly in every module. Token-economically it is neither wasteful nor ascetic.

API Cost Profile

In the Documentation Quality module, this model produces an average of 5,309 tokens against a fleet median of 3,015. That corresponds to a factor of 1.76 relative to the average across all tested models. For API usage, that is not a trivial matter: if two models document equally well but one produces nearly twice as much text, you pay the cloud provider nearly twice as much for the same air in the room.

At Anthropic’s output token prices, that is particularly relevant. Claude Sonnet 4.6 often writes thoroughly and at length, which can be useful in documentation tasks. But thoroughness tips quickly into cost when teams are generating large volumes of tickets, architecture notes, or migration guides automatically.

Code Quality and Security: Strong on Findings, Not Always Maximally Sharp on Severity

In the Code Quality module, Claude Sonnet 4.6 delivers 77.92 points. That is a strong result, and the qualitative logs explain why. In the security analysis of an intentionally vulnerable PHP application, the model identifies 18 of 19 relevant vulnerabilities with clean Markdown structure, correct terminology, and actionable fix guidance. SQL injection, plaintext passwords, path traversal, session fixation, CSRF, IDOR, type juggling, information disclosure: it lands. Above all, the prioritization of the actual attack surface lands. The model recognizes not just labels but the logic behind them.

For security-adjacent code reviews, that matters. Many models can name vulnerabilities like a vocabulary trainer. Claude Sonnet 4.6 goes a step further and explains technical relationships concisely enough to remain manageable. That makes it credible as a first reviewer — not as an auditor replacement, but as an assistant that finds more than folklore.

The picture is not without blemishes. On several critical issues, it judges a shade too conservatively. Type juggling on the API key, a sloppy admin cookie check, and an IDOR gap land only at “High” where the reference standard sets “Critical.” That is not a technical failure, but in day-to-day security work it is more than a cosmetic flaw. Organizations that label risk too politely tend to be taken seriously too late. Additionally, the log lacks an explicit attack chain from individual finding to business impact. Useful for developers, less effective for stakeholders.

On balance, Claude Sonnet 4.6 shows a very mature blend of technical accuracy and usable presentation here. On security questions it does not hallucinate wildly. It is more the type of consultant who names a vulnerability matter-of-factly and then too rarely pounds the table.

CLI and Agentic Competence: Where the Architecture Plays Its Card

The CLI benchmark result of 94.33 points is one of the clear strengths. This is exactly the zone where the agentic orientation makes sense. Claude Sonnet 4.6 thinks in steps, not in decoration. For shell-adjacent tasks, process sequences, and technical action plans, the model is very well calibrated. Such tasks benefit from the fact that it does not merely write commands, but implicitly checks what order, what caution, and what context are required.

That is worth more than it sounds at first. Many models fail in the terminal not from lack of knowledge but from lack of discipline. They produce almost-correct commands, but in a form that would be too coarse, too risky, or too imprecise on a real system. Claude Sonnet 4.6 appears markedly more controlled here. The high CLI score therefore fits very well with the curated classification as an agentic Frontier model.

However, the ToolUse result of 51.67 points shows the flip side. As soon as external tools are involved and the response must be strictly bound to the returned material, hallucinations appear. In two tool-use tasks, the model generated content that did not originate from the retrieved tool output but was fabricated. The score was consequently capped by the hallucination cap. For content-critical tasks such as research, reports, or fact-bound summaries, that is a disqualifying signal.

This is precisely where the character of this model becomes most visible. Claude Sonnet 4.6 plans well, structures well, and operates sovereignly in the technical space. But when the tool output is the only permissible reality, it occasionally allows itself literary liberties. For agent frameworks, this means: excellent as a planner and operator, riskier as a final fact editor without a verification layer.

Reasoning and Logic: Correct, Clear, Somewhat Less Deep Than the Very Best

With 79.56 points in logical reasoning, Claude Sonnet 4.6 sits clearly in the strong range. The metacog logs show a model that solves classic logic tasks cleanly, argues in well-readable structure, and makes its inference chain comprehensibly transparent. On the guard puzzle, for instance, the core logic is fully correct, the explanation is tidy, and the presentation with tables and outline is actually more accessible than some reference solutions.

What is missing is not correctness but the last measure of philosophical thoroughness. The logs note slight deficits in explicit robustness justification and in elaborating alternative formulations. Put differently: it reaches the right door, but does not stop at every lamppost along the way to label it. For most users, that is more strength than weakness.

This finding must also be read against the architecture. Although the model is classified as Thinking-capable and Thinking-Optional, the benchmark ran in the standard mode of the Anthropic API. Extended Thinking was therefore not separately activated. Given that, the result is respectable. Claude Sonnet 4.6 feels like a reasoner with the handbrake slightly on, who still drives safely. Anyone who deliberately activates adaptive Thinking in practice will likely extract somewhat more depth on more complex inferences. The benchmark deliberately evaluates the behavior a normal API user gets by default.

Documentation Quality: Competent, Thorough, Expensive in Tokens

Documentation Quality lands at 76.3 points. That is good, but not flawless, and the qualitative picture explains the number. Claude Sonnet 4.6 writes coherently, in a structured manner, and with professional gravity. It can lay out documentation in a way that teams can actually work with. No waffle, no slide poetry — generally usable text work.

The price of this solidity is length. In this area, the model produces significantly more output than the benchmark average. As long as quality rises accordingly, that is acceptable. When quality is merely good rather than outstanding, the additional text becomes a cost factor. For organizations generating thousands of documentation artifacts via the Anthropic API, that is not an academic point but budget policy in prose form.

Content Transformation and UX Writing: Strong on Tone, Vulnerable to Hard Constraints

In Content Transformation, Claude Sonnet 4.6 achieves 81.73 points; in UX Writing, 76.01. At first glance, that is a confident picture. The logs also show genuine strengths: the model can reshape language, hit tonality cleanly, and convert problematic formulations into usable, inclusive variants. In the Cultural Intelligence context, it is notable that it reliably defuses toxic or exclusionary language while remaining professional.

Yet it has a recognizable weakness when simultaneous requirements on language, length, and format are in play. This is most visible in the Content module. In one Content Transformation task, the model exceeded the explicit word limit of 250 words, producing 315 words — 126% of the limit. The system applied an automatic deduction of 20%, specifically minus 12.32 points from the achieved score. The substantive quality of the response is irrelevant at that point. The penalty applies regardless.

In a further task in the same module, the model exceeded the explicit word limit of 900 words, producing 1,409 words — 157% of the limit. Here too the system automatically applied a 20% deduction, specifically minus 17.20 points. Additionally, the model ignored the explicit language instruction and responded in English although German was required. That is not background noise but a concrete Instruction-Following weakness.

The length problem is not an isolated outlier. Across multiple Content Transformation tasks, the model shows a consistent pattern: when simultaneous constraints on language, length, and format are present, it drops the word limit as the first condition. In one video script task, the wrong output language was added to that. For production use, this is uncomfortable, because marketing, content, and communications teams in particular often work with hard format boundaries. A model that writes stylistically well but overruns the brief is like a talented columnist in form-filling: impressive, but not always helpful.

That is all the more frustrating because the actual text quality is often high. The logs for the problematic video script task credit the content with very strong production readiness, good dramaturgy, appropriate on-screen cues, and even a creative Easter egg idea. The point deduction therefore does not come from poor writing but from insufficient discipline around the guardrails. Precisely such errors hurt in benchmarks because they hurt equally in everyday use.

In the Content Transformation area, the model also ignored the explicit language instruction in one task and responded in English. In environments with a fixed target language, that is a clear deployment risk. Such errors can be corrected manually. In automated pipelines, they are poison.

Cultural Intelligence: Linguistically Sensitive, Occasionally a Touch Too Editorial

With 84.52 points, Claude Sonnet 4.6 shows one of its more pleasant sides in Cultural Intelligence. It recognizes problematic tones, removes exclusionary language, and writes with a clear sensitivity for professional, inclusive communication. That is not a minor discipline. Many real-world AI deployments revolve not around puzzles or exploits but around formulations that teams can represent externally without embarrassing themselves.

Interesting here is the nature of the error. One log praises the model for cleanly defusing a toxic job posting, but simultaneously criticizes it for overwriting the target. Rather than surgically replacing only harmful formulations, it adds its own marketing color. It occasionally turns “rewrite cleanly” into “polish slightly.” That is not a safety issue but a stylistic character trait. Claude Sonnet 4.6 is not merely an editor here but latently a co-author.

For many users, that will even be welcome. Anyone who must stay as close as possible to the input text, however, should treat the model’s first draft as a good starting point rather than a final edit.

Hallucinations: The Actual Breaking Point

The most serious qualitative warning signals concern not logic, code, or language, but hallucinations in tool use. The fact that two tool tasks were contaminated with fabricated content is not a cosmetic flaw. For agentic systems, that is precisely the red line. A model may plan, abstract, condense, and weigh. It may not, however, act as though something appeared in the tool output that was never there.

This significantly constrains the deployment recommendation. Claude Sonnet 4.6 is strong as a thinking and working model, but not blindly trustworthy as the final authority on fact-bound tool outputs. In such chains it should be combined with verification logic, structured source binding, or a downstream validation step. Anyone who omits that is building in an eloquent uncertainty factor.

Data Protection and Data Sovereignty

The data situation here is clear enough for a clean judgment. The calculated Sovereign Risk is HIGH. The reason is not a diffuse suspicion but the combination of a proprietary Anthropic model and a US provider under US law. Anthropic PBC is headquartered in San Francisco, US law including the CLOUD Act applies, and according to the Vendor Card, data is processed in the United States. For users in Germany and Europe, this means: even where contractual protective mechanisms exist, access by US authorities remains legally possible under certain conditions. That is not a theoretical culture war but applicable law.

On the positive side, a GDPR DPA is available. For organizations that must operate in GDPR compliance, that is the minimum requirement, not a bonus. Additionally, a stated data retention period of 30 days applies. That is manageable, but not zero. Anyone working with sensitive content should not let that retention period and the US-based processing disappear into the fine print. The weights provenance risk is stated as MEDIUM and aligns with the deployment reality: no open weights, full dependency on the vendor cloud.

Conclusion

Claude Sonnet 4.6 is a very strong commercial cloud model from the Anthropic API. It combines Frontier-level capability, dense architecture, agentic planning strength, Vision capability, and an enormous context window of 200,000 tokens by default and up to 1,000,000 tokens via explicit API call. In benchmark standard mode, it shows its class above all where structure, technical care, and multi-step work are required: CLI, code analysis, security review, logical reasoning. It is a model with a work ethic.

Its weaknesses are at the same time very concrete and very relevant in daily use. First: hard constraints on length and language are not always observed with the necessary discipline. Second: in tool use, it occasionally hallucinates content. That is precisely what makes the difference between “strong assistant” and “trustworthy automaton.” For coding, technical analysis, planning tasks, document structure, and agentic orchestration, Claude Sonnet 4.6 is a clear recommendation. For fact-critical research and reporting pipelines, only with a safety net. It is brilliant enough to be taken seriously, and willful enough to require oversight.

This evaluation was generated automatically based on the benchmark data. Model used: GPT-5.4 by OpenAI. The raw data and full methodology are documented in the GitHub project.