Llama 3.2 1B (Unsloth)
Llama 3.2 1B is Meta’s smallest Llama 3.2 model and a text-only baseline for edge setups. 1.23B dense parameters, 128,000 token context, locally deployable as Unsloth GGUF under the Llama 3.2 Community License. Suitable as a compact on-device reference, not as a quality anchor.
- Restricted Weights
- Nano
- llama.cpp
- Text
- Instruction-Tuned
- Restricted-Weights
- Real-Time
Sovereign Risk: LOW TODO
Key metrics
Score · Latency · Cost · Quality
- Total Score Standard
- 44.35
- Routine
- 28.17
- Reasoning
- 16.18
- LLM Judge Avg
- 1.74 / 5
- 100 Coverage
- Avg Task Duration
- 18.12s
- Real-Time
- Token Rate
- 129.8tok/s
- Output Rate
- P95 Latency
- 156.84s
- Top 5 %
- Total Tokens
- 123800
- Output Volume
- Cost per 1K
- $0
- USD / 1K Requests
- Benchmark Cost
- $0
- Total · 123800 tok
Benchmark modules
10 modules · weighted · vs. model median & top performer
Llama 3.2 1B (Unsloth)
Best model
Ø All models
Code Quality
28.1
CLI Benchmark
71.12
Logical Reasoning
38.31
UX Writing
40.05
Documentation
36.61
Content Transform.
60.06
Cultural Intelligence
49.6
Synthesis Quality
Tool Execution
ToolUse Score
Benchmark Cost
$0
Token efficiency & latency
Consumption per module vs. model median