Qwen 3 4B
Four billion parameters, Apache 2.0 license, and an optional thinking mode: Qwen 3 4B is designed for mobile applications and edge setups where every watt counts. Locally deployable under Q6 quantization, with a 128,000-token context window — generously sized for a Nano-class model. The manufacturer’s jurisdiction of China occasionally results in censored or evasive responses on politically sensitive topics.
- Open Weights
- Nano
- llama.cpp
- Text
- Real-Time
Sovereign Risk: LOW The model is run locally without a cloud connection; the CLOUD Act and data transfer risks associated with cloud usage do not apply here. Provenance remains traceable through community quantization, though the operational risk is low with purely local inference.
Key metrics
Score · Latency · Cost · Quality
- Total Score Bronze
- 60.9
- Routine
- 36.94
- Reasoning
- 23.96
- LLM Judge Avg
- 2.84 / 5
- 100 Coverage
- Avg Task Duration
- 9.95s
- Real-Time
- Token Rate
- 73.98tok/s
- Output Rate
- P95 Latency
- 24.76s
- Top 5 %
- Total Tokens
- 51400
- Output Volume
- Cost per 1K
- $0
- USD / 1K Requests
- Benchmark Cost
- $0
- Total · 51400 tok
Benchmark modules
10 modules · weighted · vs. model median & top performer
Token efficiency & latency
Consumption per module vs. model median