Swift Qwen 3.8 27B (Thinking)
Swift Qwen 3.8 27B is UkisAI’s reasoning-efficiency fine-tune on Qwen 3.8 27B: up to 58 percent fewer thinking tokens at under one percent performance loss and roughly twice the throughput on reasoning tasks. NVFP4 quantization with 262,000 tokens of context, MTP head for speculative decoding, and documented tool use — license with a commercial ARR threshold.
- Restricted Weights
- Workstation
- vLLM
- Text
- Vision
- Batch
Sovereign Risk: MEDIUM This checkpoint is a UkisAI fine-tune and an NVFP4 quantization of Qwen/Qwen3.8-27B. The base lineage is documented, but the weights are distributed under the gated Swift Open License v1.0 with an ARR threshold and are optimized for Blackwell/vLLM deployment — provenance is therefore clear, but not fully open in the OSS sense.
Key metrics
Score · Latency · Cost · Quality
- Total Score Gold
- 80.48
- Routine
- 49.15
- Reasoning
- 31.33
- LLM Judge Avg
- 4.09 / 5
- 100 Coverage
- Avg Task Duration
- 110.61s
- Batch
- Token Rate
- 15.96tok/s
- Output Rate
- P95 Latency
- 290.93s
- Top 5 %
- Total Tokens
- 115800
- Output Volume
- Cost per 1K
- $0
- USD / 1K Requests
- Benchmark Cost
- $0
- Total · 115800 tok
Benchmark modules
10 modules · weighted · vs. model median & top performer
Token efficiency & latency
Consumption per module vs. model median