Qwen 3.8 Omni Flash

One million tokens of context for text, image, audio, and video in a single model: Qwen3.8-Omni-Flash (available September 18, 2026) is Alibaba’s first omni-modal model with an agentic focus — it understands multi-hour audio and video material, plans tasks, and executes them via tool calling. At 0.15 / 0.47 USD per million tokens, Alibaba claims it reduces audio costs by more than 98 percent compared to its predecessor. Output is text only; the weights remain closed.

Alibaba Version 3.8 Commercial use permitted Dense 1000 K Context $0.15 / $0.47 per 1M

  • Proprietary
  • Frontier
  • OpenRouter
  • Text
  • Vision
  • Unusable

Sovereign Risk: MEDIUM-HIGH The model is developed by a Chinese company and deployed exclusively via cloud APIs (weights are not released). As a Chinese company, Alibaba is subject to PRC data security and cybersecurity laws, not the US CLOUD Act. Since no weights are distributed, there is no redistribution risk, but there is a potential data access risk by Chinese authorities for data processed through Alibaba Cloud regions.

Key metrics

Score · Latency · Cost · Quality

Total Score Silver
75.27
Routine
45.9
Reasoning
29.37

Rank #33

LLM Judge Avg
3.76
100 Coverage
Avg Task Duration
165.38
Unusable
Token Rate
42.07
Output Rate
P95 Latency
479.21
Top 5 %
Total Tokens
410500
Output Volume
Cost per 1K
$0.0005
USD / 1K Requests
Benchmark Cost
$0.19
Total · 410500 tok

Benchmark modules

10 modules · weighted · vs. model median & top performer

Qwen 3.8 Omni Flash Best model Ø All models
Code Quality 81.76
CLI Benchmark 85
Logical Reasoning 68.95
UX Writing 73.31
Documentation 71.68
Content Transform. 70.81
Cultural Intelligence 77.84
Synthesis Quality 66.67
Tool Execution 90
ToolUse Score 77.67
Benchmark Cost $0.19

Token efficiency & latency

Consumption per module vs. model median

Token consumption per module

Performance profile