Qwen 3.6 Plus

Qwen 3.6 Plus is Alibaba’s proprietary flagship model of the Qwen 3.6 series, featuring a hybrid MoE architecture with a focus on agentic coding and multimodal processing. With a one-million-token context window, configurable thinking mode, and native agentic capabilities, the model targets demanding production applications. Available exclusively via cloud APIs; Chinese jurisdiction applies.

Alibaba Version 3.6 Plus Commercial use permitted MoE 1000 K Context 02/2026 $0.325 / $1.95 per 1M

  • Proprietary
  • Frontier
  • OpenRouter
  • Text
  • Vision
  • Video
  • Instruction-Tuned
  • Agentic Orchestrator
  • Batch

Sovereign Risk: HIGH The model is operated exclusively via the Alibaba Cloud API. Data transmitted through the API is subject to China’s National Security Law (NSL), which may enable state access to data. Local deployment is not possible — no weight download is available.

Tool-use profile: 6 assets in detail

Comparison of asset performance (P1/P2/Combined) against the fleet average

Asset performance (radar)

Score Breakdown vs. Fleet Average


Tool-use details

Asset performance, reliability, and runtime profile

CrucibleMark evaluates tool use across 6 independent tests. Click a test name for details.

Reliability

  • Tool Call Valid: No
  • Retry: Not required
  • Hallucination: Not detected

Reliability measures how consistently a model actually executes tool calls: Tool Call Valid schema and format accepted, Retry Required successful only after retry, Hallucination Flag fabricated tools or parameters detected. All three green means production-ready.

Runtime profile

Call 1
6.54
First Request
MCP
0.68
Protocol Latency
Synthesis
34.23
Response Generation
Total
248.76
Sum of All Phases
Token
22861
Input + Output
Cost
$0
Cost per Run

The runtime profile shows the latency and cost metrics for the model run: Call 1 First Request, MCP Protocol Latency, Synthesis Response Generation, Total sum of all phases, supplemented by Token input and output and Cost cost per run.

Tool-use review

Updated · Instruction-Tuned · Agentic Orchestrator

Deployment Verdict

Conditional deploy, because tool execution is mostly strong, but synthesis fidelity remains too unreliable and the tool call during the run was not consistently valid.

Tool Execution Profile

Qwen 3.6 Plus demonstrates genuine tool intelligence rather than mere template usage. On the Web Search & Tool Selection test — which checks the choice between search and direct retrieval without an explicit hint — it reliably identifies the need for web_search. On the URL Construction test, which measures the derivation of a target URL from internal knowledge followed by a fetch, it performs adequately but less deterministically. This points to flexible planning rather than rigid pattern-following.

For an MCP pipeline, however, the picture is not entirely clean. P1 is a solid 68.33, but the signal “tool call valid: false” is relevant. In practice, this means the model understands the flow most of the time but does not produce protocol-clean calls at every step. Since no retry was required, the issue lies more in execution precision than in a fundamental misunderstanding.

Synthesis Fidelity

How well does it condense tool results? Only with limited reliability. P2 sits at 50.00 and variance is high. HTTP Fetch & Extract is still decent, Tool Failure Handling (404) and URL Construction & Fetch are good, but EU License Research falls off noticeably in condensation quality. Most critically, Multilingual Search & Synthesis: quality drops sharply when cross-lingual research requires a German-language summary. For pipelines that need concise, dependable result condensation, this is too inconsistent.

Does it stay with the tool result or fall back on training data? In the Honeypot EU License Research test — which checks whether current license restrictions are answered from web sources rather than training knowledge — it does not hallucinate. That is the most important trust signal. The P2 score of 40, however, shows that it does not cleanly convert the retrieved material into a reliable answer. Trust in provenance is therefore stronger than trust in condensation.

Error Resilience

On the 404 test, which distinguishes transparent failure from fabricated fallback content, Qwen 3.6 Plus responds in a production-appropriate manner. It communicates the failure rather than inventing page content. This is acceptable for real tool chains and considerably more important than stylistic response quality.

Operational Profile

Total 248.76s per run. Call 1 6.54s, Call 2 34.23s, MCP latency 0.68s. Slow for the performance shown. Price: $0.325 per 1M input tokens, $1.95 per 1M output tokens. Cost-efficient to moderate on the API side, but runtime undermines the economics.

Conclusion & Recommendation

Suitable for agentic research pipelines with human review, particularly where tool selection and transparent error handling matter more than perfect final condensation. Not suitable for compliance, policy, or multilingual synthesis pipelines where the final answer is processed downstream without review. Additionally, the cloud-only operating model under Chinese jurisdiction is an independent disqualifying factor for sensitive tool data.

This analysis was generated automatically based on benchmark data. Model used: GPT 4.5 by OpenAI. The raw data and full methodology are documented in the GitHub project.