Qwen 3.5 397B A17B

Qwen 3.5 397B A17B is Alibaba’s first natively multimodal Open Weights Frontier model, processing text, image, and video within a single model. Its hybrid architecture combining Gated DeltaNet and Sparse MoE activates only 17 billion of the total 397 billion parameters per token, with a context window of 256,000 tokens. Fully commercially usable under the Apache 2.0 license.

Alibaba Version 3.5 Commercial use permitted MoE 397 B (17 B active) 262 K Context 12/2025 $0.39 / $2.34 per 1M

  • Open Weights
  • Frontier
  • OpenRouter
  • Text
  • Vision
  • Instruction-Tuned
  • Agentic Orchestrator
  • Batch

Sovereign Risk: MEDIUM Alibaba Cloud is a Chinese company and subject to China’s National Security Law (NSL). When using the Alibaba Cloud API, government access to transmitted data is possible. Purely local inference with public Open Weights significantly reduces this risk — the CLOUD Act analogue (NSL) is only directly relevant when using the cloud API.

Tool-use profile: 6 assets in detail

Comparison of asset performance (P1/P2/Combined) against the fleet average

Asset performance (radar)

Score Breakdown vs. Fleet Average


Tool-use details

Asset performance, reliability, and runtime profile

CrucibleMark evaluates tool use across 6 independent tests. Click a test name for details.

Reliability

  • Tool Call Valid: No
  • Retry: Not required
  • Hallucination: Not detected

Reliability measures how consistently a model actually executes tool calls: Tool Call Valid schema and format accepted, Retry Required successful only after retry, Hallucination Flag fabricated tools or parameters detected. All three green means production-ready.

Runtime profile

Call 1
3.19
First Request
MCP
0.86
Protocol Latency
Synthesis
29.63
Response Generation
Total
202.03
Sum of All Phases
Token
5442
Input + Output
Cost
$0.0052
Cost per Run

The runtime profile shows the latency and cost metrics for the model run: Call 1 First Request, MCP Protocol Latency, Synthesis Response Generation, Total sum of all phases, supplemented by Token input and output and Cost cost per run.

Tool-use review

Created · Instruction-Tuned · Agentic Orchestrator

Deployment Verdict

Conditional deploy, because tool execution appears reliable and no hallucination was detected, but synthesis quality at Combined 78.80 is not stable enough for high-stakes result consolidation.

Tool Execution Profile

Qwen 3.5 397B A17B behaves in a production-near manner at the execution level. Tool calls were valid, MCP-compliant, and required no retries. This is the most important baseline safeguard for a tool pipeline. In the Web Search & Tool Selection test, which checks — without an explicit hint — whether the model chooses between search and direct fetch, the model reliably identifies the appropriate tool class. This argues against a rigid call pattern and in favor of genuine tool selection. In the URL construction test, it constructs the target URL usably, but not precisely enough for deterministic pipelines with tight error tolerance. Execution is therefore strong, but not blindly trustworthy when the path must be derived from model knowledge.

An important note for context: this model is primarily a vision-language system. The text-tool competence visible here is therefore credible, but does not reflect its full product surface.

Synthesis Fidelity

How well does it consolidate tool results? Solid, but not at Frontier level. P2 of 68 shows that it often merges retrieved information correctly, but loses precision in doing so. This is particularly evident in the Multilingual Search & Synthesis test, which evaluates cross-language research and German-language summarization: the search succeeds, but the consolidation falls noticeably short of the execution quality.

Does it stay within the tool result or fall back on training data? There is no data on this from the Honeypot EU License Research. The only positive signal is indirect: no hallucination was detected in the available runs. For compliance or licensing pipelines this is helpful, but it is not a substitute for a passed honeypot.

Error Resilience

In the 404 test, which checks whether a failed tool call is handled transparently or whether fabricated page content appears, the model stays on the safe side. It does not hallucinate despite the error. This is acceptable for production. The P2 of 60 indicates, however, that error communication is functional but not always optimally consolidated or guided.

Operational Profile

Total 190.54s per run. Call 1: 3.13s. MCP latency: 0.79s. Call 2: 34.19s. Slow. Cost per run: 0.004944. Inexpensive to very inexpensive for this size class. Price-to-performance is good; latency remains the operational bottleneck.

Conclusion & Recommendation

Suitable for MCP pipelines where clean tool execution matters more than perfect final consolidation: research agents, discovery workflows, multimodal preprocessing stages, and assisted analyst tooling. Not the first choice for compliance, multilingual executive summaries, or other paths where the response itself is the product. If you deploy it, do so with downstream validation of summaries and clear guards for URL derivation and final user-facing text.

This analysis was generated automatically based on benchmark data. Model used: GPT 4.5 by OpenAI. The raw data and full methodology are documented in the GitHub project.