Llama 3.2 1B (Unsloth)

Llama 3.2 1B is Meta’s smallest Llama 3.2 model and a text-only baseline for edge setups. 1.23B dense parameters, 128,000 token context, locally deployable as Unsloth GGUF under the Llama 3.2 Community License. Suitable as a compact on-device reference, not as a quality anchor.

Meta Version 3.2 Commercial use permitted Dense 1.23 B (1.23 B active) 128 K Context 12/2023 locally tested

  • Restricted Weights
  • Nano
  • llama.cpp
  • Text
  • Instruction-Tuned
  • Restricted-Weights
  • Real-Time

Sovereign Risk: LOW TODO

Key metrics

Score · Latency · Cost · Quality

Total Score Standard
44.35
Routine
28.17
Reasoning
16.18

Rank #99

LLM Judge Avg
1.74
100 Coverage
Avg Task Duration
18.12
Real-Time
Token Rate
129.8
Output Rate
P95 Latency
156.84
Top 5 %
Total Tokens
123800
Output Volume
Cost per 1K
$0
USD / 1K Requests
Benchmark Cost
$0
Total · 123800 tok

Benchmark modules

10 modules · weighted · vs. model median & top performer

Llama 3.2 1B (Unsloth) Best model Ø All models
Code Quality 28.1
CLI Benchmark 71.12
Logical Reasoning 38.31
UX Writing 40.05
Documentation 36.61
Content Transform. 60.06
Cultural Intelligence 49.6
Synthesis Quality
Tool Execution
ToolUse Score
Benchmark Cost $0

Token efficiency & latency

Consumption per module vs. model median

Token consumption per module

Performance profile