Qwen 3.8 27B (Thinking)

Qwen3.8-27B is Alibaba’s dense 27.8-billion-parameter variant of the Qwen3.8 family (August 14, 2026), natively multimodal with a vision encoder for text, image, and video under the Apache 2.0 license. The 64-layer hybrid architecture of Gated DeltaNet plus Gated Attention blocks plus Multi-Token Prediction delivers 262,144 tokens of context (extensible to approximately one million via YaRN), configurable reasoning (xhigh/medium/low), and an optional no-thinking mode.

Alibaba Version 3.8 Commercial use permitted Dense 27.8 B 262 K Context 12/2025 locally tested

  • Open Weights
  • Workstation
  • vLLM
  • Text
  • Vision
  • Video
  • Instruction-Tuned
  • Long Context
  • Batch

Sovereign Risk: MEDIUM The model was developed by Alibaba, a company headquartered in China. This carries a theoretical risk due to Chinese legislation. However, since the weights have been released under the permissive Apache 2.0 license and are intended for local deployment, no data is transmitted to the manufacturer. The risk is rated ‘medium’: the origin lies in a high-risk jurisdiction, but the open license and local usage substantially minimize the practical risk.

Tool-use profile: 6 assets in detail

Comparison of asset performance (P1/P2/Combined) against the fleet average

Asset performance (radar)

Score Breakdown vs. Fleet Average


Tool-use details

Asset performance, reliability, and runtime profile

CrucibleMark evaluates tool use across 6 independent tests. Click a test name for details.

Reliability

  • Tool Call Valid: No
  • Retry: Not required
  • Hallucination: Detected

Reliability measures how consistently a model actually executes tool calls: Tool Call Valid schema and format accepted, Retry Required successful only after retry, Hallucination Flag fabricated tools or parameters detected. All three green means production-ready.

Runtime profile

Call 1
5.95
First Request
MCP
1.45
Protocol Latency
Synthesis
41.48
Response Generation
Total
293.29
Sum of All Phases
Token
14558
Input + Output
Cost
$0
Cost per Run

The runtime profile shows the latency and cost metrics for the model run: Call 1 First Request, MCP Protocol Latency, Synthesis Response Generation, Total sum of all phases, supplemented by Token input and output and Cost cost per run.

Tool-use review

Created · Instruction-Tuned · Long Context

Deployment Verdict

Conditional deploy: tool usage is strong, but one invalid tool call and detected hallucination limit confidence for unsupervised MCP pipelines. The overall impression is good, but not robust enough for high-trust workloads.

Tool Execution Profile

Qwen 3.8 27B demonstrates genuine tool intelligence rather than mere pattern-following. In the Web Search & Tool Selection test — which requires choosing between search and direct retrieval without an explicit hint — it reliably selects the appropriate tool. This points to workable planning logic in dynamic pipelines. In the URL Construction test, which requires deriving the target URL from internal knowledge and then fetching it correctly, it is less precise. The result is usable, but not deterministic enough for systems that generate requests directly from model output.

A P1 score of 90 reflects strong execution overall. The critical issue, however, is that the tool call in that run was marked invalid. This is not a retry issue — not a mere formatting problem followed by a clean correction — but a reliability signal: the model can make good tool decisions, yet does not consistently produce MCP-compliant calls.

Synthesis Fidelity

How well does it consolidate tool results? Only moderately. The P2 score of 59.17 fits the profile: solid on clear error cases and structured search, but weaker on precise extraction and consolidation from fetched content. HTTP Fetch & Extract — clean uptake of concrete facts from retrieved pages — drops noticeably, with a P2 of 35. For production pipelines, this means results often need to be cross-checked or re-normalized after the tool call.

Does it stay within the tool result or fall back on training data? In the Honeypot EU License Research test — which checks whether current licensing restrictions are answered from web sources rather than training knowledge — it stays on the safe side. No hallucination was detected there. At the same time, global hallucination detected is set to true. This should be treated as a security risk, not merely a quality issue. Once a model outputs fabricated facts as a tool result, the entire tool infrastructure becomes vulnerable.

Error Resilience

In the 404 test — which checks for transparent behavior when a retrieval fails — the model responds in a production-appropriate manner. It communicates the error openly and does not fabricate page content. This is a strong signal. Such behavior is acceptable for operational pipelines because failures remain visible and downstream systems can escalate cleanly.

Operational Profile

Call 1: 5.95s. MCP latency: 1.45s. Call 2: 41.48s. Total: 293.29s.
Local: no API costs.
Speed: slow to very slow across the full run, relative to only good overall performance.

Conclusion & Recommendation

Suitable for locally operated research and orchestration pipelines with human oversight, especially where tool selection matters more than perfect consolidation. Also usable for 404-resilient agent paths. Not suitable for compliance, fact-checking, or extraction workloads where tool results are passed downstream without modification. Anyone deploying Qwen 3.8 27B should enforce strict output validation, schema checking, and a second verification stage after every tool step.

This analysis was generated automatically based on benchmark data. Model used: GPT-5.4 by OpenAI. The raw data and full methodology are documented in the GitHub project.