Qwen3.8-2.4T-A95B

Qwen3.8-2.4T-A95B is Alibaba’s first Open Weights release at Qwen-Max level (August 12, 2026), a fine-grained Mixture-of-Experts model with 2.4 trillion total and 95 billion active parameters per token. License: proprietary ‘Qwen3.8-Max License’. The model processes text in a 262,144-token context (expandable to approximately one million) with a mandatory reasoning mode (low/high/xhigh) and hybrid attention combining Gated-DeltaNet and Gated-Attention.

Qwen Version 3.8 Commercial use permitted MoE 2400 B (95 B active) 262 K Context $2 / $6 per 1M

  • Open Weights
  • Frontier
  • OpenRouter
  • Text
  • Long Context
  • Agentic Orchestrator
  • Batch

Sovereign Risk: HIGH The model is developed by the Qwen Team at Alibaba Cloud, a company headquartered in China. Due to Chinese legislation (including the National Security Law) and the associated potential for state influence over technology companies, the origin risk of the weights is classified as high, regardless of the deployment location.

LLM Model Review

Created on · Long Context · Agentic Orchestrator

With an overall score of 67.23%, Qwen3.8-2.4T-A95B is a model with audible ambition and visibly unsteady execution. As an agentically oriented Frontier generalist with mandatory reasoning, 95 billion active parameters in a MoE architecture, and a 262K context window, it aims to plan, structure, and carry long workloads. In the benchmark, it manages this only with friction: strong on CLI, decent on Content and Culture, but conspicuously weak where precision under length and format pressure should count. The Cloud Open Weights variant was tested via OpenRouter in the provider’s default mode — without a separate thinking toggle; the speed profile badge reads Batch Tool Expert and aptly describes the model as a tool for thorough, batch-oriented rather than interactive work. Sovereign Risk: HIGH — developer and provider context sit with Alibaba Cloud in China; Chinese data and security laws therefore apply, with clear sovereignty risks for European users in cloud deployments.

Header Notes: Stability and Reliability

Metric Value Rating Analysis
Timeout Rate 10/49 Unreliable The model is unreliable and drops out significantly often in practice. For a Cloud Open Weights model via OpenRouter, this is not an abstract lab finding but a real API risk: endpoint scatter, overload, or network issues hit the user directly.
P95 Response Time 234.5 s Critical Extreme tail latency. The model scatters massively and is unsuitable for time-critical processes. In five percent of cases, the wait is long enough that assistance quickly becomes standstill.

Architecture and Character

The pre-assigned category fits this model’s profile surprisingly well. Qwen3.8-2.4T-A95B is not a sharply trimmed direct-answer system but a large, planning-oriented system with a tendency toward internal preparation. This is visible in several places. First, it is an Agentic Orchestrator: it often decomposes tasks sensibly, visibly thinks in steps, and delivers structure rather than elegance on complex, multi-part prompts. Second, it is a Thinking-Mandatory model. Reasoning here is not a switch but a character trait. This is precisely why answers are frequently well-conceived but not infrequently weighed down by their own effort.

The right benchmark matters here. The total parameter count of 2.4 trillion sounds like heavy industry but is secondary in practice. What is relevant is the 95 billion active parameters per token. For a Frontier MoE model, that is still enormous — but not the raw full-activation that the total figure implies. The hybrid attention combining Gated-DeltaNet and Gated-Attention targets long contexts and efficiency. In theory, this is the right tool for extended work contexts, document chains, and agentic planning. In the benchmark, what shows up most is the tendency toward long reasoning paths. The intelligence is there. The discipline is not always.

Performance and Speed Profile

The Batch Tool Expert badge is more than a cosmetic label. It describes a model not built for rapid back-and-forth but for workbench tasks with tool and structure involvement. Qualitatively, this means moderately to sluggishly paced responses, especially on difficult tasks, but with ambition toward more complete reasoning. For a mandatorily thinking agentic model, this is not inherently a flaw. Such models often pay for internal planning with time.

Important context: Qwen3.8-2.4T-A95B ran here as a Cloud Open Weights model via OpenRouter. The measured generation speed is therefore primarily a finding about the provisioned cloud infrastructure plus network path, not about any hypothetical self-hosted installation. Especially with Open Weights in the cloud: speed here is a provider characteristic. And this provider picture is not impressive. The long tail and high failure rate turn the batch character into not a virtue but often a test of patience.

Reasoning and Logic

In the reasoning module, Qwen3.8-2.4T-A95B shows the side for which one considers such models in the first place. The guard-and-doors test is solved correctly, explained cleanly, and presented in good German. Particularly noteworthy: the model does not merely name the correct question but also briefly sorts and dismisses competing approaches. This is exactly what one expects from a mandatorily thinking system.

The quality is not spectacularly brilliant, but reliably good. The golden standard lacks some didactic polish — clearer visualization or more tightly structured reasoning blocks, for instance. But the core is sound: logical correctness, traceable derivation, no smoke and mirrors. For an agentically oriented model, this is the right priority. It prefers to think methodically rather than brilliantly.

CLI, Tool Proximity, and Security-Adjacent Tasks

With 84.34 points in the CLI domain and 90.0 on Tool Execution, the tooling layer is clearly among this model’s stronger zones. This is consistent. Agentic Orchestrator models are not primarily built to pull the perfect one-liner out of thin air but to arrange work steps into a usable sequence. Precisely this strategic tool proximity is present here.

The model also shows substance in security-adjacent reasoning. In the Code Quality protocol, it correctly identifies central vulnerabilities: SQL injection, plaintext passwords, path traversal, CSRF gaps, session issues, weak reset tokens, and several implicit vulnerabilities. Particularly important: the five hidden security problems are explicitly captured as such. That is not a minor checkbox item but the difference between superficial and serious analysis.

Unfortunately, the model does not reliably follow its good thinking through to completion.

Code Quality: Strong Analysis, Weak Delivery

Code Quality is the domain where Qwen3.8-2.4T-A95B reveals its character most clearly. Substantively, it understands the task. Formally, it fails itself too often. The benchmark average of 61.8 points is simply too thin for a Frontier model of this ambition.

The judge’s security verdict on a PHP audit task is almost frustratingly positive in a negative sense: the model correctly identifies the major and hidden vulnerabilities, categorizes them cleanly, and formulates concise, usable fix hints. The Markdown table holds. The German holds. Then the response breaks off mid-table. Not because the model goes off the rails substantively, but because its internal reasoning consumes the output budget. That is the difference between a sharp analyst and one who, at the decisive moment, pulls the report halfway out of the typewriter.

In the Code Quality domain, one output breaks off mid-table — the response is technically truncated, not a content error. The score deduction results from the incomplete response, not from substantive shortcomings.

A structural pattern is also at work here. In three Code Quality tasks, internal reasoning tokens displaced the output budget so far that only incomplete results remained. In one task, the model was left with a residual budget of just 1,682 tokens after massive internal reasoning; in two others, 1,291 and 968 tokens respectively. The effect is practically the same: the analysis starts strong, the delivery ends too early. For a model intended for planning and complex security analysis, this is not a minor concern. Anyone writing security audits cannot regularly leave the table with a truncated table.

UX Writing: The Major Weakness

At 53.27 points, UX Writing is the most visible breaking point. Not because the model fundamentally lacks command of language, but because it loses composure under simultaneous constraints of tone, idiom, inclusion, and brevity.

The qualitative protocol on the job posting rewrite illustrates the pattern clearly and unpleasantly. Qwen3.8-2.4T-A95B meets the formal requirements, removes toxic language, and applies gender-neutral phrasing. But the text sounds mechanical. The “ninja” pathos is not transformed into an elegant rewrite but into a rather dutifully managed neutralization. Phrases like “convince through quality in competition” feel stilted. Inclusion is appended as an explicit afterthought, whereas the better solution weaves it unobtrusively into the text. The result is usable but not refined. An HR professional would let it pass. It does not inspire enthusiasm.

The systemic problem of the reasoning budget compounds this. In two UX Writing tasks, the model’s internal reasoning consumption damaged its output budget. In one task, 2,543 tokens remained visible. In another, the remainder fell to exactly zero. That is the brutal case: not a poor formulation, but no complete formulation at all. For UX work, this is devastating, because nuance here is not a bonus but the product.

Documentation Quality: Labored to Brittle

Documentation Quality lands at 46.12 points — a range where a Frontier model has no excuses. Long contexts and structured explanatory work should be precisely where Qwen3.8-2.4T-A95B is at home. The architecture practically cries out for documentation, summarization, and layered explanation. All the more striking, then, how often the internal reasoning depth damages external usability.

Three Documentation Quality tasks suffered from the same problem: reasoning tokens displaced the actual output budget. Once 1,510 tokens remained, once 941, once 1,614. This sounds technical but translates plainly in practice: the model spends too much time and budget preparing, then delivers the documentation only in truncated form. For product documents, migration guides, or operational runbooks, this is a real risk. Anyone who receives a guide cut off mid-way has not achieved half a victory but, in all likelihood, an expensive mistake.

Content Transformation: Strong Craft, Long Shadow

Content Transformation at 77.83 points is one of the clear strengths. This is deserved. The YouTube tutorial protocol shows a model that not only meets complex production requirements but casts them into a near-broadcast-ready script: analysis at the specified length, complete timestamps, screen cues, retention elements, an Easter egg, traceable why-explanations. This is not a lucky hit but solid transformation craft.

Notably, the model maintains structure even under high detail load. The one small timing deviation on the Pattern Interrupt is manageable. Here, Qwen3.8-2.4T-A95B works as one would want from an agentically shaped long-context model: it keeps the production framework in mind and does not lose track of secondary conditions.

But even here, the reasoning mode is not without consequences. In one Content Transformation task, internal reasoning consumption constrained the output budget. The score deduction therefore does not stem from substantive failure but from this model’s structural problem: it thinks lavishly and charges the bill to the visible output.

Cultural Intelligence: Clean, but Not Finely Chiseled

Cultural Intelligence at 77.84 points is convincing. Here, Qwen3.8-2.4T-A95B demonstrates that it is not blunt in linguistic and cultural terms. It reliably maintains the German target language, recognizes problematic terms, and replaces them in accordance with the rules. In detail, it occasionally lacks the idiomatic elegance of the best responses, but it does not tip into embarrassing awkwardness.

This matters especially for German-language business texts. The model is not a great stylist. It is more the correct, conscientious editor who sees tonal problems but does not always find the most elegant turn of phrase. For many corporate contexts, that is sufficient. For brand-defining communication, less so.

API Cost Profile

Qwen3.8-2.4T-A95B is a Cloud Open Weights model. Its output volume is therefore not merely a stylistic question but directly a cost question. And here the model behaves as though output were free.

In the CLI domain, it produces an average of 5,876 tokens against a fleet median of 314. That is 18.71 times the average across all tested models. In Code Quality, it generates 13,868 tokens against a fleet median of 2,906 — a factor of 4.77. In UX Writing, 9,679 tokens versus a median of 1,511, which is 6.41 times the median. Documentation Quality at 11,679 versus 2,966 and Content Transformation at 8,013 versus 1,790 also sit massively above the field.

This would be easier to forgive if quality rose proportionally. It does not. In several modules, the model produces considerably more text and still fails on completeness or precision. For API deployment, this means very concretely: higher costs at often only average added value. At the listed prices of $2 per million input tokens and $6 per million output tokens, this is not a theoretical cosmetic flaw but a budget issue with advance notice.

Data Privacy and Data Sovereignty

For European organizations, this model is not an incidental detail from a data protection standpoint but a red file. The calculated Sovereign Risk is HIGH. The reasoning is doubly clear: the weights originate from the Qwen Team at Alibaba Cloud in China, and the provider context is subject to Chinese law — specifically PIPL, CSL, and DSL.

China is listed as the data location. A verified data processing agreement under GDPR is not known; the DPA status is listed as unknown. The data retention period is likewise not reliably documented and stands at -1 days — effectively unknown. For organizations required to operate in GDPR compliance, this is a serious compliance obstacle. Without a clear DPA and without an EU adequacy decision, technical suitability does not translate into approved enterprise deployment.

The distinction between deployment and provenance also matters. Even if the model is accessed through a different cloud provider, the weights provenance risk remains high, because the origin of the weights lies with Alibaba in China. This is not a political footnote but part of the risk profile.

Conclusion

Qwen3.8-2.4T-A95B is an idiosyncratic, large reasoning model with genuine substance and a tendency toward self-sabotage. It can plan, it can identify security problems, it can shape complex transformation tasks into remarkably well-formed outputs. Its CLI and tool proximity is strong, its Cultural and Content profile decent to good. But it falls too hard on Documentation Quality and especially UX Writing, and the combination of high verbosity, critical tail latency, and ten failures in 49 tests is simply too noisy to ignore in productive pipelines.

For agentic batch workflows with human review, the model is interesting. For interactive assistance, reliable documentation pipelines, or linguistically refined end-user texts, it is too erratic in this form. The core problem is not insufficient intelligence but insufficient housekeeping: too much internal reasoning, too little reliable delivery. Across all tests, no noteworthy hallucinations — Qwen3.8-2.4T-A95B fails more on overreach and abrupt cutoff than on fabrication.

This evaluation was generated automatically based on the benchmark data. Model used: GPT 4.5 by OpenAI. The raw data and full methodology are documented in the GitHub project.