Qwen 3.8 Omni Flash
One million tokens of context for text, image, audio, and video in a single model: Qwen3.8-Omni-Flash (available September 18, 2026) is Alibaba’s first omni-modal model with an agentic focus — it understands multi-hour audio and video material, plans tasks, and executes them via tool calling. At 0.15 / 0.47 USD per million tokens, Alibaba claims it reduces audio costs by more than 98 percent compared to its predecessor. Output is text only; the weights remain closed.
- Proprietary
- Frontier
- OpenRouter
- Text
- Vision
- Unusable
Sovereign Risk: MEDIUM-HIGH The model is developed by a Chinese company and deployed exclusively via cloud APIs (weights are not released). As a Chinese company, Alibaba is subject to PRC data security and cybersecurity laws, not the US CLOUD Act. Since no weights are distributed, there is no redistribution risk, but there is a potential data access risk by Chinese authorities for data processed through Alibaba Cloud regions.
Key metrics
Score · Latency · Cost · Quality
- Total Score Silver
- 75.27
- Routine
- 45.9
- Reasoning
- 29.37
- LLM Judge Avg
- 3.76 / 5
- 100 Coverage
- Avg Task Duration
- 165.38s
- Unusable
- Token Rate
- 42.07tok/s
- Output Rate
- P95 Latency
- 479.21s
- Top 5 %
- Total Tokens
- 410500
- Output Volume
- Cost per 1K
- $0.0005
- USD / 1K Requests
- Benchmark Cost
- $0.19
- Total · 410500 tok
Benchmark modules
10 modules · weighted · vs. model median & top performer
Token efficiency & latency
Consumption per module vs. model median