Qwen 3.8 Flash-Next (NVIDIA) (Thinking)
As an open preview of the Qwen4 architecture, Alibaba introduces Qwen3.8-Flash-Next — an experimental MoE model that NVIDIA has prepared as an NVFP4 quantization for local inference. Of approximately 180 billion parameters on disk, only around 6 billion activate per token, with text, image, and video input and 262,000 tokens of native context. License: combined NVIDIA and Qwen license. Important: This preview is not the hosted API Qwen3.8-Flash.
- Open Weights
- Server
- vLLM
- Text
- Vision
- Video
- Instruction-Tuned
- Long Context
- Interactive
Sovereign Risk: MEDIUM The base model originates from Alibaba (CN); the NVFP4 distribution by NVIDIA (US) reduces operational risk somewhat, but results in ‘medium’ due to US jurisdiction (e.g., CLOUD Act). Additional identity risk: Qwen3.8-Flash-Next is explicitly an experimental Open Weights preview of the upcoming Qwen4 architecture, separate from the production-hosted ‘Qwen3.8-Flash’ API with more production features.[374][383][384]
Key metrics
Score · Latency · Cost · Quality
- Total Score Gold
- 81.68
- Routine
- 49.6
- Reasoning
- 32.08
- LLM Judge Avg
- 4.07 / 5
- 100 Coverage
- Avg Task Duration
- 39.13s
- Interactive
- Token Rate
- 42.72tok/s
- Output Rate
- P95 Latency
- 97.8s
- Top 5 %
- Total Tokens
- 112600
- Output Volume
- Cost per 1K
- $0
- USD / 1K Requests
- Benchmark Cost
- $0
- Total · 112600 tok
Benchmark modules
10 modules · weighted · vs. model median & top performer
Token efficiency & latency
Consumption per module vs. model median