Qwen3.8-2.4T-A95B

Qwen3.8-2.4T-A95B is Alibaba’s first Open Weights release at Qwen-Max level (August 12, 2026), a fine-grained Mixture-of-Experts model with 2.4 trillion total and 95 billion active parameters per token. License: proprietary ‘Qwen3.8-Max License’. The model processes text in a 262,144-token context (expandable to approximately one million) with a mandatory reasoning mode (low/high/xhigh) and hybrid attention combining Gated-DeltaNet and Gated-Attention.

Alibaba Version 3.8 Commercial use permitted MoE 2400 B (95 B active) 262 K Context $2 / $6 per 1M

  • Open Weights
  • Frontier
  • OpenRouter
  • Text
  • Long Context
  • Agentic Orchestrator
  • Unusable

Sovereign Risk: HIGH The model is developed by the Qwen Team at Alibaba Cloud, a company headquartered in China. Due to Chinese legislation (including the National Security Law) and the associated potential for state influence over technology companies, the origin risk of the weights is classified as high, regardless of the deployment location.

Tool-use profile: 6 assets in detail

Comparison of asset performance (P1/P2/Combined) against the fleet average

Asset performance (radar)

Score Breakdown vs. Fleet Average


Tool-use details

Asset performance, reliability, and runtime profile

CrucibleMark evaluates tool use across 6 independent tests. Click a test name for details.

Reliability

  • Tool Call Valid: No
  • Retry: Not required
  • Hallucination: Not detected

Reliability measures how consistently a model actually executes tool calls: Tool Call Valid schema and format accepted, Retry Required successful only after retry, Hallucination Flag fabricated tools or parameters detected. All three green means production-ready.

Runtime profile

Call 1
4.59
First Request
MCP
1.55
Protocol Latency
Synthesis
25.2
Response Generation
Total
188.06
Sum of All Phases
Token
25494
Input + Output
Cost
$0
Cost per Run

The runtime profile shows the latency and cost metrics for the model run: Call 1 First Request, MCP Protocol Latency, Synthesis Response Generation, Total sum of all phases, supplemented by Token input and output and Cost cost per run.

Tool-use review

Created · Long Context · Agentic Orchestrator

Deployment Verdict

Conditional deploy. The model is strong at tool execution and does not hallucinate in this run, but the invalid tool call at an overall merely good yield makes it acceptable for production MCP pipelines only with guardrails in place.

Tool Execution Profile

Qwen3.8-2.4T-A95B demonstrates genuine tool intelligence rather than mere schema-following. In the Web Search & Tool Selection test — which checks the choice between search and direct retrieval without an explicit hint — it correctly identifies that web_search is needed before fetch. That is a good signal for open, dynamic agent paths. In the URL Construction test, which requires deriving the correct target URL from internal knowledge followed by a subsequent fetch, it is serviceable but not deterministic enough. This is where the distinction lies: strategic tool selection is strong; operational precision in the concrete call varies. Since the tool call was not valid overall and no retry was required, this points more toward an execution or format edge case than a fundamental misunderstanding of the task.

Synthesis Fidelity

How well does it condense tool results? Only adequately. The P2 score of 66.67 aligns with the individual results: HTTP Fetch & Extract is very clean, Multilingual Search & Synthesis drops off noticeably. The model can structure retrieved information but loses precision when condensing multilingual or compliance-adjacent content. For architectures in which the tool layer supplies only raw material and the model builds the final report, this is a limiting factor.

Does it stay within the tool result or fall back on training? In the Honeypot EU License Research test — which checks whether current license restrictions are actually retrieved from web sources — it remains in the safe zone overall. No hallucination finding is the more important signal here than the merely average P2 yield. Trust in the tool chain therefore holds, even if the synthesis is not consistently reliable.

Error Resilience

Acceptable for production. In the Tool Failure Handling (404) test, which targets transparency when tool calls fail, the model communicates the error rather than fabricating page content. That is exactly what a robust pipeline requires. The finding is not excellent, but it is safe.

Operational Profile

Call 1: 4.59s. MCP latency: 1.55s. Call 2: 25.20s. Total: 188.06s. Slow for the quality level achieved. Cost/run: local. Model price profile: $2.0/1M input, $6.0/1M output. Not expensive for Frontier-level, but runtime is clearly the operational bottleneck.

Conclusion & Recommendation

Suitable for agentic research and orchestration pipelines in which tool selection, long contexts, and cautious error handling matter more than perfect final synthesis. Not the first choice for compliance reports, multilingual synthesis, or strictly deterministic MCP pipelines where every tool call must be formally correct. Deploy only with call validation, output schema checks, and a downstream verification layer for the final summary.

This analysis was generated automatically based on benchmark data. Model used: GPT-5.4 by OpenAI. The raw data and full methodology are documented in the GitHub project.