Qwen 3.8 Flash-Next (NVIDIA) (Thinking)

As an open preview of the Qwen4 architecture, Alibaba introduces Qwen3.8-Flash-Next — an experimental MoE model that NVIDIA has prepared as an NVFP4 quantization for local inference. Of approximately 180 billion parameters on disk, only around 6 billion activate per token, with text, image, and video input and 262,000 tokens of native context. License: combined NVIDIA and Qwen license. Important: This preview is not the hosted API Qwen3.8-Flash.

NVIDIA Version 3.8-Next Commercial use permitted MoE 180 B (6 B active) 262 K Context locally tested

  • Open Weights
  • Server
  • vLLM
  • Text
  • Vision
  • Video
  • Instruction-Tuned
  • Long Context
  • Interactive

Sovereign Risk: MEDIUM The base model originates from Alibaba (CN); the NVFP4 distribution by NVIDIA (US) reduces operational risk somewhat, but results in ‘medium’ due to US jurisdiction (e.g., CLOUD Act). Additional identity risk: Qwen3.8-Flash-Next is explicitly an experimental Open Weights preview of the upcoming Qwen4 architecture, separate from the production-hosted ‘Qwen3.8-Flash’ API with more production features.[374][383][384]

Key metrics

Score · Latency · Cost · Quality

Total Score Gold
81.68
Routine
49.6
Reasoning
32.08

Rank #2

LLM Judge Avg
4.07
100 Coverage
Avg Task Duration
39.13
Interactive
Token Rate
42.72
Output Rate
P95 Latency
97.8
Top 5 %
Total Tokens
112600
Output Volume
Cost per 1K
$0
USD / 1K Requests
Benchmark Cost
$0
Total · 112600 tok

Benchmark modules

10 modules · weighted · vs. model median & top performer

Qwen 3.8 Flash-Next (NVIDIA) (Thinking) Best model Ø All models
Code Quality 82.04
CLI Benchmark 88.33
Logical Reasoning 76.62
UX Writing 73.55
Documentation 85.98
Content Transform. 85.02
Cultural Intelligence 82.36
Synthesis Quality 62.5
Tool Execution 90
ToolUse Score 82.83
Benchmark Cost $0

Token efficiency & latency

Consumption per module vs. model median

Token consumption per module

Performance profile