Claude Opus 5

Claude Opus 5 is Anthropic’s Opus flagship as of July 24, 2026, positioned as an everyday model between Opus 4.8 and the more expensive Fable 5. The cloud-only model under US jurisdiction offers 1 million tokens of context, 128,000 tokens of output, and adaptive reasoning control with five effort levels (low/medium/high/xhigh/max). Mid-conversation tool switching without cache loss and a Fast Mode with 2.5× speed round out the offering.

Anthropic Version 5 Commercial use permitted Dense 1000 K Context 05/2026 $5 / $25 per 1M

  • Proprietary
  • Frontier
  • Anthropic
  • Text
  • Vision
  • Agentic Orchestrator
  • Long Context
  • Interactive

Sovereign Risk: MEDIUM Anthropic is a US-based company and subject to the CLOUD Act. The model weights are proprietary and not distributed; no additional risk from weight distribution. Data handling is governed by Anthropic Commercial Terms.

LLM Model Review

Updated on · Agentic Orchestrator · Long Context

Claude Opus 5 achieves an overall score of 78.73 percent and carries the Speed Profile Badge Interactive Doc Writer on the Leaderboard. That fits the character of this model remarkably well: less of a go-getter for frantic tool chains, more of a high-caliber knowledge worker with a sense for structure, depth, and extended context. As a commercial cloud model via the Anthropic API, it clearly plays in the Frontier class: agentic use case, dense parameter architecture, 1-million-token context, plus factory-default behavior with no switchable thinking mode during the test run. Sovereign Risk: HIGH — Anthropic, as a US company, is subject to the CLOUD Act; data is processed in the United States.

Header Notes: Stability and Reliability

Metric Value Rating Analysis
Timeout Rate 3/49 Sporadic The model shows sporadic failures that would require retries in practice.
P95 Response Time 117.67 s Problematic Significant outliers that interrupt workflow.

For a Frontier model from the vendor’s own cloud, these header notes are not a minor footnote — they are a deployment signal. Three failures in 49 tests are not a total breakdown, but enough to breed distrust in agent workflows or longer production chains. Add to that a heavy tail in response times: anyone embedding Claude Opus 5 in interactive processes gets not only quality, but occasionally a waiting-room atmosphere as well.

Architecture and Character: What Claude Opus 5 Is Built For

The pre-assigned architecture category hits the mark with surprising precision. Claude Opus 5 is not a generic all-purpose chatbot but an agentic orchestration model with reasoning DNA, a very large context window, and a multimodal design. In practice, this means: it thinks in structures, not in isolated bursts. It plans, organizes, weighs options, and holds long threads together. Anyone expecting exact shell one-liners with the mechanical precision of a specialized tool from a model like this is measuring against the wrong use case. Anyone who instead needs analysis, synthesis, documentation, and multi-step problem decomposition is much closer to its actual strengths.

The evaluation framework matters here. Claude Opus 5 is editorially classified as Use Case agentic, Size Class Frontier, and Parameter Architecture dense. No excuses for small model size, no leniency for a specialty niche, no MoE bonus for clever activation economics. Frontier in a dense configuration means: the highest expectations apply. That is exactly the standard against which Claude Opus 5 must be measured.

The run was conducted under the endpoint’s factory-default behavior; no switchable thinking mode exists here. That is not a side note. For Claude, “default” does not mean simple. It only means that the reasoning architecture operates internally, without the user configuring it separately in the benchmark. The result is typically Claude: often deep, often deliberate, occasionally a little too enamored with its own intelligence.

Performance, Speed, and Cost Profile

The Interactive Doc Writer badge is more than decoration. It signals a model suited for high-quality, interactive knowledge work: documentation, structured long-form content, demanding transformations, explanatory responses. That is precisely where Claude Opus 5 delivers its most convincing moments. The generation speed does not feel fast in the sense of “snappy” — it feels more cultivated. The text arrives with care. For a thinking and agentic model, that is not automatically a flaw, but it is a cost center in real-time scenarios.

On pricing, the picture is clear: $5 per 1 million input tokens, $25 per 1 million output tokens. This is not an impulse purchase but a professional tool with a professional invoice. Because Claude Opus 5 works in a token-disciplined manner overall, the cost lands less painfully than one might fear from a deeply reasoning model. It does not exceed the expected verbosity envelope in any module. The model behaves with token economy. For a system that is visibly tuned toward analysis over action, that is a remarkable form of self-restraint.

Reasoning and Logic: Strong in Thinking, Weaker in Deference to Format

In the reasoning module, Claude Opus 5 demonstrates why the Thinking category is not merely a label here. The substantive logic holds. On the classic guard puzzle, the model delivers the correct core solution, explains it thoroughly, dismantles naive approaches, and even adds alternative formulations. This is not superficial pattern-matching but genuine working-through of the problem. You can sense it: this model does not just want to answer — it wants to cleanly illuminate the structure of the problem.

The price for that is a familiar Claude issue: formal discipline does not always come first. In the metacognition protocol, the explicitly required <thought> tags were absent, as was a clearly separated final answer. Substantively, this is strong. Formally, it is non-compliant. For humans, this often still reads as authoritative; for automated pipelines, it is simply a risk. An agent that thinks brilliantly internally but ignores the template is like a very intelligent employee who improves every form — including, unfortunately, the ones where you needed exactly that form.

Metacognition Compliance (Reasoning): The model refuses to use the explicitly requested <thought> tags in 3/5 metacog tests, citing a consistent policy statement. The reasoning content itself is correct — the score deduction results from format non-compliance, not from errors in thinking. For comparison: in the tag-free reasoning_5* tests, the model achieves an average score of approximately 77.24 percent, which is on par with other strong models. CrucibleMark deliberately evaluates native zero-shot instruction compliance as a real-world characteristic — this deduction is methodologically intentional.

This is the decisive point: Claude Opus 5 does not fail here primarily on logic, but on compliance. Those using it for analysis, evaluation, or conceptual sparring get substance. Those using it for strictly formatted output in machine-readable pipelines should plan for guardrails, validators, or post-formatting.

Code Quality and Security: Analytically Strong, but Not Maximally Aggressive

In the code and security domain, Claude Opus 5 delivers a good to very good performance. The audit rating of 81.92 percent feels earned. In the protocols, the model identifies virtually the entire relevant vulnerability surface: SQL injection in multiple variants, plaintext passwords, type juggling, IDOR, path traversal, XSS, weak token generation, mail header injection, session fixation. This is not erratic name-dropping but technically precise identification of attack surfaces with actionable remediation hints.

Its real strength lies in the organization of findings. Claude Opus 5 sorts security vulnerabilities readably, prioritizes comprehensibly, and formulates countermeasures in a way that not only auditors but also development teams can work with. For many real-world security reviews, that is more valuable than a fireworks display of obscure edge cases. Where the model falls short of a perfect security analysis is in the exploit narrative. The judge protocols rightly note that concrete attack chains and proof-of-concept narratives are less sharply articulated than the reference standard. Claude names the mine. It does not always paint the crater.

For a model classified as agentic Frontier, this is interesting. It analyzes excellently, orchestrates plausibly, and delivers solid remediation proposals. But it operates less like an offensive red-teamer and more like a very experienced security architect who does not need to frighten the board with spectacular attack paths to be right. For defensive work, that is often exactly right. For offensive security work, there are sharper tools.

CLI and Tool Execution: Where the Orchestrator Heritage Becomes Both an Excuse and a Limit

The weakest visible domain is tool execution. With 83.33 percent in the ToolUse score and 80.94 percent in the CLI domain, Claude Opus 5 does not fall behind — but it does not shine there as effortlessly as it does in documentation or security analysis. That is not surprising. The Agentic-Orchestrator category is precisely the right context for this: such models are built to decompose subtasks and delegate to specialized tools or subagents when needed. In practice, that is often smarter than pressing out the last exact flag in a one-liner yourself.

Still, a limit remains. The benchmark measures out-of-the-box behavior. And there, Claude Opus 5 shows that it does not always have the hardness of a specialized DevOps model when it comes to direct format and execution precision. That is forgivable, but not without consequence. Those looking for a central planning model for agent systems will find substance here. Those wanting a model that delivers script-exact output with maximum mechanical precision without a second look should orient themselves toward other endpoints.

Content Transformation and Documentation: Where Claude Opus 5 Plays Its Class

Perhaps the most convincing side of Claude Opus 5 emerges in tasks that demand more than correctness. In the Documentation Quality domain, the model reaches 86.01 percent. In the Content Transformation module, it achieves 81.25 percent. These are not mere effort scores. This is the signature of a model that shapes complex raw material into usable form.

The video script for 2FA setup is a good example. Claude Opus 5 precisely identifies the gaps in the template, adds a hook, timing, production notes, troubleshooting, retention mechanics, and even a workable Easter egg. The judge certifies the result as production-ready. Tasks like this clearly suit the model: establishing structure, making implicit requirements explicit, combining tone with utility. There, Claude does not feel like a text generator but like a remarkably clear-eyed editor with production awareness.

The model is also strong in classical documentation work. The “Interactive Doc Writer” badge is not marketing fiction — it captures the benchmark character quite accurately. Claude Opus 5 does not merely write a lot. It organizes knowledge. That is a distinction that matters in practice.

UX Writing, Cultural Intelligence, and Instruction-Following: Linguistically Strong, Not Always Obedient

Language quality is consistently high. Particularly in German, Claude Opus 5 delivers clean, natural, and professional text. In the Cultural Intelligence domain, however, it trips over a classic Claude banana peel: it can solve the task linguistically with authority and still fail on a simple directive.

The HR rewriting protocol is instructive here. The German text itself was well-crafted, inclusive, and professional. The problem was not the quality of the rewrite, but the decision to append extensive English explanations even though this was explicitly prohibited. This is not a careless slip. It is a compliance defect. Claude Opus 5 wants to be helpful, even when “helpful” would have meant simply keeping quiet for once.

This pattern runs subtly through multiple modules. The model is rarely grossly off-target. It violates instructions more from excess than from inability: too much explanation, too much framing, too little ascetic adherence. For humans, this is often endearing. For production systems with hard output requirements, it is a structural weak point.

Hallucinations and Content Reliability

Substantively, Claude Opus 5 indulges in very little fabrication. It reads, across most of its output, like a model that prefers to analyze and delimit rather than paper over gaps with elegant invention. This fits the overall signature of this system: high structural fidelity, strong knowledge organization, a cautious rather than imaginative response policy. In security and documentation contexts, that is worth its weight in gold, because any invented certainty there quickly becomes expensive.

Data Privacy and Data Sovereignty

For European organizations, Claude Opus 5 is not neutral ground from a data protection standpoint — it is a deliberate risk decision. The provider is Anthropic PBC, headquartered in San Francisco; the applicable law is US (CLOUD Act); the data location per the vendor card is the USA; data retention is 30 days. A GDPR DPA is available, which for many organizations creates the minimum contractual basis for a GDPR-compliant path. That does not fundamentally defuse the situation, however.

The calculated Sovereign Risk is HIGH. The rationale is clear: US jurisdiction without EU-level safeguards. The CLOUD Act means that US authorities can, under certain conditions, demand access to data — even where individual technical or organizational protective measures are in place. For German and European organizations, this is not a theoretical footnote but part of the risk calculus. The Weights Provenance Risk is MEDIUM: no additional uncertainty from redistributed weights, but a fundamental dependency on a proprietary US vendor.

Conclusion

Claude Opus 5 is a very good Frontier model with a clearly recognizable professional ethos: think, organize, explain, structure. It is strong in documentation, strong in content transformation, strong in security analysis, and substantively convincing in reasoning. It weakens where blind obedience to format specifications matters more than intellectual quality. That is precisely what defines its character. This model is not a willing stenographer but a self-assured senior consultant. That impresses — until you need exact output formats.

For deployment, this means: highly recommended for knowledge work, review processes, long contexts, security and quality analysis, and as a central reasoning model in agentic systems. Less ideal for strictly formalized tool output without safeguards, for latency-critical interaction, and for environments where every retry causes operational pain. Across all tests, no notable hallucinations — the model prefers to invent little rather than embarrass itself with false certainty.

Within the Claude family, Opus 5 reads like the serious documentarian among the large models: cultivated, substantive, occasionally over-eager to explain, and with a pace you have to appreciate. Those looking for a brilliant cloud model that holds long threads cleanly in hand will find a great deal of quality here. Those who need absolute precision under time pressure should not send Claude Opus 5 onto the stage alone.

This evaluation was generated automatically based on the benchmark data. Model used: GPT-5.4 by OpenAI. The raw data and full methodology are documented in the GitHub project.