Qwen 3.8 Omni Flash

One million tokens of context for text, image, audio, and video in a single model: Qwen3.8-Omni-Flash (available September 18, 2026) is Alibaba’s first omni-modal model with an agentic focus — it understands multi-hour audio and video material, plans tasks, and executes them via tool calling. At 0.15 / 0.47 USD per million tokens, Alibaba claims it reduces audio costs by more than 98 percent compared to its predecessor. Output is text only; the weights remain closed.

Alibaba Version 3.8 Commercial use permitted Dense 1000 K Context $0.15 / $0.47 per 1M

  • Proprietary
  • Frontier
  • OpenRouter
  • Text
  • Vision
  • Unusable

Sovereign Risk: MEDIUM-HIGH The model is developed by a Chinese company and deployed exclusively via cloud APIs (weights are not released). As a Chinese company, Alibaba is subject to PRC data security and cybersecurity laws, not the US CLOUD Act. Since no weights are distributed, there is no redistribution risk, but there is a potential data access risk by Chinese authorities for data processed through Alibaba Cloud regions.

Tool-use profile: 6 assets in detail

Comparison of asset performance (P1/P2/Combined) against the fleet average

Asset performance (radar)

Score Breakdown vs. Fleet Average


Tool-use details

Asset performance, reliability, and runtime profile

CrucibleMark evaluates tool use across 6 independent tests. Click a test name for details.

Reliability

  • Tool Call Valid: No
  • Retry: Not required
  • Hallucination: Not detected

Reliability measures how consistently a model actually executes tool calls: Tool Call Valid schema and format accepted, Retry Required successful only after retry, Hallucination Flag fabricated tools or parameters detected. All three green means production-ready.

Runtime profile

Call 1
5.8
First Request
MCP
1.55
Protocol Latency
Synthesis
83.76
Response Generation
Total
546.63
Sum of All Phases
Token
27794
Input + Output
Cost
$0
Cost per Run

The runtime profile shows the latency and cost metrics for the model run: Call 1 First Request, MCP Protocol Latency, Synthesis Response Generation, Total sum of all phases, supplemented by Token input and output and Cost cost per run.

Tool-use review

Created

Deployment Verdict

Conditional deploy, because overall tool usage is strong, but call validity is not clean enough and synthesis sits at only a mid-production level. A combined 77.67 is viable, but not self-sustaining without guardrails.

Tool Execution Profile

Qwen 3.8 Omni Flash demonstrates genuine tool intelligence, not just rigid pattern behavior. In the Web Search & Tool Selection test, it correctly recognizes — without an explicit hint — that current information requires a search rather than a direct fetch. That is a strong signal for agentic orchestration. It also performs reliably on the tool side in EU License Research and Multilingual Search & Synthesis.

Formal execution is weaker. Tool call valid is overall false, even though P1 sits high at 90. This does not point to a comprehension problem, but to sloppiness in individual calls or parameters. In the URL Construction test — which checks whether the model can derive the target address itself and then retrieve it — it reaches only usable rather than deterministic precision. The HTTP Fetch & Extract test shows the same pattern: access is mostly correct, but not robust enough for pipelines that depend on strictly reproducible calls.

Synthesis Fidelity

How well does it consolidate tool results? Only solidly. A P2 of 66.67 means the model can merge results but loses precision when condensing facts closely. Quality drops noticeably in particular in the HTTP Fetch & Extract test, which measures exact extraction from real page content. For reports and first drafts, that is sufficient. For compliance, contract, or regulatory summaries, it is too loose.

Does it stay within the tool result or fall back on training? Mostly yes, and that is the most important positive finding. In the honeypot EU License Research — which checks whether current license restrictions genuinely come from web sources — it does not hallucinate. P2 60 is not strong on substance, but acceptable on the trust side: it does not fabricate seemingly current facts outside the tool basis.

Error Resilience

In the 404 test, the model responds in a production-appropriate manner. It communicates the failure transparently and does not hallucinate substitute content. P2 80 in this scenario is a good signal for safe degradation: the pipeline remains auditable when a tool fails.

Operational Profile

Call 1: 5.80s. Call 2: 83.76s. MCP latency: 1.55s. Total per run: 546.63s. Clearly slow. Cost: $0.15/1M input, $0.47/1M output. Inexpensive for a frontier model, but the runtime is heavy relative to the synthesis performance demonstrated.

Conclusion & Recommendation

Suitable for MCP pipelines involving research, tool selection, and controlled error handling — especially when current web data matters more than perfect consolidation. Not suitable as the final authority for precise extraction, regulatory summaries, or strictly deterministic fetch flows without additional validation. Deploy only with schema checks, response validation, and a downstream review of extracted facts.

This analysis was generated automatically based on benchmark data. Model used: GPT-5.4 by OpenAI. The raw data and full methodology are documented in the GitHub project.