Qwen 3 4B

Four billion parameters, Apache 2.0 license, and an optional thinking mode: Qwen 3 4B is designed for mobile applications and edge setups where every watt counts. Locally deployable under Q6 quantization, with a 128,000-token context window — generously sized for a Nano-class model. The manufacturer’s jurisdiction of China occasionally results in censored or evasive responses on politically sensitive topics.

Alibaba Version 3 Commercial use permitted Dense 4 B (4 B active) 128 K Context 09/2024 locally tested

  • Open Weights
  • Nano
  • llama.cpp
  • Text
  • Real-Time

Sovereign Risk: LOW The model is run locally without a cloud connection; the CLOUD Act and data transfer risks associated with cloud usage do not apply here. Provenance remains traceable through community quantization, though the operational risk is low with purely local inference.

Key metrics

Score · Latency · Cost · Quality

Total Score Bronze
60.9
Routine
36.94
Reasoning
23.96

Rank #88

LLM Judge Avg
2.84
100 Coverage
Avg Task Duration
9.95
Real-Time
Token Rate
73.98
Output Rate
P95 Latency
24.76
Top 5 %
Total Tokens
51400
Output Volume
Cost per 1K
$0
USD / 1K Requests
Benchmark Cost
$0
Total · 51400 tok

Benchmark modules

10 modules · weighted · vs. model median & top performer

Qwen 3 4B Best model Ø All models
Code Quality 57.1
CLI Benchmark 73.34
Logical Reasoning 57.95
UX Writing 58.65
Documentation 62.77
Content Transform. 64.77
Cultural Intelligence 48.3
Synthesis Quality 51.67
Tool Execution 90
ToolUse Score 70.5
Benchmark Cost $0

Token efficiency & latency

Consumption per module vs. model median

Token consumption per module

Performance profile