Llama 3.2 3B (Unsloth)

3.21B dense parameters, 128,000 tokens context: Llama 3.2 3B is Meta’s compact text-only variant of the Llama 3.2 family for local tasks such as summarization, paraphrasing, and instruction-following. Unsloth GGUF build, Llama 3.2 Community License, fully operable offline.

Meta Version 3.2 Commercial use permitted Dense 3.21 B (3.21 B active) 128 K Context 12/2023 locally tested

  • Restricted Weights
  • Nano
  • llama.cpp
  • Text
  • Instruction-Tuned
  • Restricted-Weights
  • Real-Time

Sovereign Risk: LOW TODO

Key metrics

Score · Latency · Cost · Quality

Total Score Bronze
52.05
Routine
32.2
Reasoning
19.86

Rank #98

LLM Judge Avg
2.37
100 Coverage
Avg Task Duration
8.46
Real-Time
Token Rate
55.02
Output Rate
P95 Latency
17.18
Top 5 %
Total Tokens
52200
Output Volume
Cost per 1K
$0
USD / 1K Requests
Benchmark Cost
$0
Total · 52200 tok

Benchmark modules

10 modules · weighted · vs. model median & top performer

Llama 3.2 3B (Unsloth) Best model Ø All models
Code Quality 46.8
CLI Benchmark 70.01
Logical Reasoning 45.05
UX Writing 54.15
Documentation 43.29
Content Transform. 60.97
Cultural Intelligence 57.3
Synthesis Quality 31.67
Tool Execution 63.33
ToolUse Score 47.83
Benchmark Cost $0

Token efficiency & latency

Consumption per module vs. model median

Token consumption per module

Performance profile