Qwen3.8-Flash
Qwen3.8 Flash is Alibaba’s cloud-only multimodal reasoning model with a one-million-token context window and pricing of $0.16 / $0.47 per million tokens. It processes text, image, and video with tool calling, accessible via Alibaba Cloud and OpenRouter. The open base Qwen3.8-Flash-Next is available, but the tested build is cloud-only. Architecture and parameter count are not disclosed, and data is routed through Chinese jurisdiction.
- Open Weights
- Frontier
- OpenRouter
- Text
- Vision
- Video
- Agentic Orchestrator
- Long Context
- Interactive
Sovereign Risk: MEDIUM Base weights are openly available (Qwen/Qwen3.8-Flash-Next, qwen-community-1.0, official NVFP4/GGUF/FP8 quants) — local deployment possible. The tested OpenRouter endpoint, however, is Alibaba’s production build (Qwen3.8-Flash, 1M context, built-in tools), whose exact post-training differences from Flash-Next have not been disclosed. Development/operations are subject to Chinese jurisdiction.
Key metrics
Score · Latency · Cost · Quality
- Total Score Silver
- 79.51
- Routine
- 48.19
- Reasoning
- 31.31
- LLM Judge Avg
- 4.02 / 5
- 100 Coverage
- Avg Task Duration
- 39.68s
- Interactive
- Token Rate
- 50.94tok/s
- Output Rate
- P95 Latency
- 88.89s
- Top 5 %
- Total Tokens
- 113600
- Output Volume
- Cost per 1K
- $0.0005
- USD / 1K Requests
- Benchmark Cost
- $0.05
- Total · 113600 tok
Benchmark modules
10 modules · weighted · vs. model median & top performer
Token efficiency & latency
Consumption per module vs. model median