Llama 3.2 3B (Unsloth)

3.21B dense parameters, 128,000 tokens context: Llama 3.2 3B is Meta’s compact text-only variant of the Llama 3.2 family for local tasks such as summarization, paraphrasing, and instruction-following. Unsloth GGUF build, Llama 3.2 Community License, fully operable offline.

Meta Version 3.2 Commercial use permitted Dense 3.21 B (3.21 B active) 128 K Context 12/2023 locally tested

  • Restricted Weights
  • Nano
  • llama.cpp
  • Text
  • Instruction-Tuned
  • Restricted-Weights
  • Real-Time

Sovereign Risk: LOW TODO

LLM Model Review

Created on · Instruction-Tuned · Restricted-Weights

With an overall score of 52.05%, Llama 3.2 3B (Unsloth) remains firmly in the zone of serviceable lightweight work — but not sovereign overall performance. The Speed Profile Badge “Real-Time DevOps Expert” signals high responsiveness, and indeed the model feels nimble rather than sluggish at everyday pace. The content picture is considerably rougher: for a Nano generalist model with 3.21 billion dense parameters and an instruct character, directness is expected — but the pre-assigned Thinking classification reads more like a promise on the packaging than lived practice in this particular run, especially since the test ran explicitly in Standard Mode with Thinking disabled.

Header Grades: Stability and Reliability

Metric Value Rating Analysis
Timeout Rate 0/49 Stable The model ran with absolute stability and reliability throughout testing.
P95 Response Time 17.18 s Consistent Very low tail latency, virtually no outliers.

For a local Nano model, the header grades are almost the most important good news. Anyone deploying small models wants a tool, not a diva. That is exactly what Llama 3.2 3B (Unsloth) delivers here: no dropouts, no stalls, no meltdown under load. This meaningfully raises its practical value, even if it obviously does nothing to argue away the quality deficits.

Architecture, Ambition, and What You Can Fairly Ask of This Model

The curated classification here is more than a label. Llama 3.2 3B (Unsloth) is a Generalist — not a specialized code tool, not a pure reasoning engine, and not a multimodal system. It belongs to the Nano class: the weight class where you can expect efficiency, local usability, and workable instruction-following, but not deep world knowledge or reliable mastery of multi-step tasks. Add to that a Dense architecture: all 3.21 billion parameters are active on every response. There is no MoE trick here inflating the apparent model size. What you get is exactly the capacity listed on the spec sheet.

The Thinking and Instruct tags create an interesting tension. Instruct fits the observed behavior well: the model responds compactly, often rule-oriented, sometimes too schematically. Thinking, however, must be immediately qualified in this report. This run took place in Standard Mode — with Thinking disabled. Shorter, direct answers are not a flaw in this setup; they are the defined operating state. The problem is not that Llama answers briefly here. The problem is that even within that brevity, it too often stops halfway on logic, prioritization, and accuracy of detail.

Speed and Token Economy

As a local model, Llama 3.2 3B (Unsloth) ran on an NVIDIA DGX Spark (GB10 Grace Blackwell Superchip, ~115 GB Unified Memory — no practical memory limit for tested model sizes). The “Real-Time DevOps Expert” badge fits the character of this run: the model generates visibly fast enough for interactive use, quick follow-up queries, and simple editorial or developer dialogues. It is not a batch worker for long thinking sessions — more the nimble desk assistant.

Token economy matters here. Llama 3.2 3B (Unsloth) behaves in a consistently token-economical manner. No module exceeds the expected verbosity range. On the contrary: across all measured modules it stays below the fleet median, sometimes considerably so. For a local model, this is not merely an aesthetic value — it is directly a practical one. Less text here generally means less waiting time. The flip side is visible, however: particularly in Code Quality, Documentation, and Reasoning, the brevity does not feel elegantly condensed — it feels substantively starved. This model is not just saving words. It is too often saving the critical details as well.

Code Quality: Cleanly Formatted, Technically Thin

The Code Quality score of 46.8 is not an operational accident — it is an accurate summary of the model’s character. Llama 3.2 3B (Unsloth) can sketch security issues, format tables, and name obvious vulnerabilities. It recognizes SQL Injection, XSS, and some more advanced topics. For 3.21 billion parameters, that is not nothing. But one should not project depth onto this surface.

The qualitative log makes the weakness very clear. In a security analysis, the model lists 9 vulnerabilities while the reference standard identifies 19 — meaning it misses more than half the problem spectrum. More seriously: it misclassifies several findings. A client-side-manipulable admin flag is downgraded, type juggling on the API key is misread in its criticality, plaintext passwords and path traversal are rated too leniently. These are not cosmetic errors. In security practice, exactly these kinds of misjudgments cause teams to close the wrong issues first.

The fixes also stay at form-filling level. “Prepared statements,” “sanitization,” “admin flag management” — this sounds like security, but is often just the label on the toolbox, not the repair itself. Concrete countermeasures, attack chains, or defensible prioritization are absent. Anyone trying to harden real code with this needs a second instance with greater technical reach. In short: the model sees that the house is on fire. It just does not reliably say which floor, or what to use to put it out.

Reasoning and Logic: Visibly Trying, Not Reliable

The Logical Reasoning score of 45.05 is weak, and the logs explain why. On a classic guard riddle, the model delivers a formally well-structured response in German and even uses the required <thought> tags. That matters, because it shows: there is no compliance or formatting problem here. The problem is the logic itself.

Instead of the well-known, robust meta-question directed at one of the guards, the model essentially proposes: “If I open the other door, will I die?” This is not a clever alternative — it is simply the wrong question. It provides no guaranteed disambiguation between the liar and the truth-teller. The model projects confidence without cleanly working through both cases. That is precisely where genuine reasoning separates from plausible-sounding improvisation.

For the assigned Thinking category, this is sobering — even accounting fairly for Standard Mode. The benchmark does not penalize the model for being brief. It penalizes the brevity for not holding up. A good small model does not need to reason at length. It does need to reason correctly on core questions. Llama 3.2 3B (Unsloth) too often delivers the tone of certainty without the structural integrity of the argument. That is the most dangerous form of error, because it looks competent at first glance.

Content Transformation: Workable Structure, Weak Editorial Quality

In Content Transformation, the model reaches 60.97 — one of its better scores. This fits the profile: summarizing, rephrasing, rough structuring — these tasks suit compact instruct models more than deep analysis. But here too, the result should not be mistaken for professionalism.

The video script log shows a typical Nano compromise. Llama 3.2 3B (Unsloth) understands the task in principle. It delivers analysis, timestamps, production notes, a hook, a CTA, and even an Easter egg. But these elements often feel like mandatory fields in a form, not components of a compelling narrative arc. The Judge criticizes contradictory time markers, generic direction notes, formal rather than spoken language, and an emotional flatness that is fatal for video formats. A sentence that reads like a user manual does not suddenly speak to people just because you put a timecode in front of it.

There is also a structural length problem. In one Content Transformation task, the model exceeded the explicit word limit of 250 words, delivering 369 words148% of the limit. The system applied an automatic deduction of 20%, or -11.84 points, to the achieved score. The substantive quality of the response is therefore irrelevant — the penalty applies regardless. The word limit here is not a decorative suggestion; it is part of the task.

This length problem is not an isolated outlier. Together with the UX Writing area, the model shows a consistent pattern: when facing simultaneous constraints on language, format, and length, it drops the word limit as the first condition to go. For Nano models, this is not unusual. But typical does not mean harmless. In editorial and agent workflows, this kind of failure is immediate.

UX Writing: Competent Craft with a Tendency to Overrun

The UX Writing score of 54.15 is middling, but less bleak than some other disciplines. The model can structure microcopy, build tables, and set up tasks in a formally orderly way. The delivered fragments show that it at least recognizes step sequences and Progressive Disclosure. For short interface texts and first drafts, that is genuinely usable.

But here too, the hard reality of constraints applies. In one UX Writing task, the model exceeded the explicit word limit of 350 words, delivering 442 words126% of the limit. The system applied an automatic deduction of 20%, or -11.00 points. The same rule applies: miss the length constraint and you fail, even if the text would otherwise be acceptable. For product teams this matters, because microcopy operates on exactly the opposite principle of “a little more is fine.” A dialog box has no patience.

This behavior can be read as a weakness of small instruct models — and that would not be wrong. They often prioritize semantic task completion over precise simultaneous adherence to multiple side constraints. But this explanation does not help the user much. If text has to fit a UI, a missed word limit is not a philosophical problem — it is a broken ticket.

Documentation Quality: Willing to Inform, but Not Solid

The score of 43.29 in Documentation Quality is among the weaker areas, and that is not surprising. Good technical documentation demands not just language competence, but also hierarchization, precision, avoidance of gaps, and the ability to meet the reader at exactly the right point. That combination is hard for Nano Generalists.

The log fragments point to a recurring pattern: the model understands roughly what needs to be documented, but not reliably how deep, how operational, or how robust it needs to be. It frequently delivers a workable shell. What is missing is the final layer of completeness, prioritization, and concrete actionability. For internal notes or a first rough draft, that may suffice. For publication-ready technical documentation, it does not.

Cultural Intelligence: Good Intentions, Poor Execution

At 57.3, Cultural Intelligence fares better than Reasoning or Documentation — but the qualitative finding is harsher than the number suggests. Particularly revealing is the German-language task, in which problematic language from a job posting was to be rewritten in inclusive form. The model refuses to engage entirely, responding in effect that it cannot assist with discriminatory or degrading content.

This is a classic misstep of over-cautious safety patterns. The task was specifically to remove the toxic language. The model confuses remediation with propagation of the problem. For productive editorial work, this is frustrating, because it blocks legitimate redaction work. The real irony is bitter: the model does not fail out of malice, but out of moral short-circuit logic. Good intentions, zero utility.

Tool Execution, Security, and Hallucinations: This Is Where It Gets Serious

The Tool Use score of 31.67 is weak, and the automated violation summary makes unmistakably clear why. In two Tool Use tasks, the model hallucinated content that did not originate from the retrieved tool result but was fabricated. The system capped the P2 score via the hallucination cap. For content-critical tasks such as research, factual reports, or tool-assisted evaluation, this is a disqualifying signal.

Data Privacy and Data Sovereignty

A dedicated privacy section is not necessary here, because Llama 3.2 3B (Unsloth) ran locally with its own weights and did not go through a cloud provider. On provenance, the finding is straightforwardly favorable: the weights provenance risk is rated LOW; the weights originate from the Meta Llama lineage in an Unsloth build. The relevant caveat lies less in the data-leakage question than in the license nature: the Llama 3.2 Community License permits commercial use, but is explicitly Restricted Weights and not OSI-open.

Conclusion

Llama 3.2 3B (Unsloth) is a small local model with clearly recognizable utility and equally clear limits. It responds quickly, runs stably on the test system, and works token-efficiently. For simple rephrasing, short helper texts, quick drafts, and uncomplicated assistance tasks, it is a perfectly reasonable candidate. Anyone looking for a Nano model for offline-capable everyday work gets no miracle here — but a tool.

As soon as tasks demand multiple simultaneous constraints, however, the picture shifts. Logic is fragile, security judgments are incomplete, culturally sensitive transformations can tip into refusal, and on tool-bound factual tasks the model hallucinates in ways that should not be explained away. Add to that repeated word-limit violations with automatic score penalties. That is not a pedantic side issue — it is genuine workflow damage.

On balance, Llama 3.2 3B (Unsloth) is a typical Nano Generalist with an agreeable runtime character and limited cognitive reach. Fine for Edge-adjacent, local assistance. Too risky for security analyses, reliable reasoning, editorially sensitive rewrites, or tool-critical production pipelines. This model works quickly and politely. It just does not work deeply enough, often enough.

This evaluation was generated automatically based on the benchmark data. Model used: GPT-5.4 by OpenAI. The raw data and full methodology are documented in the GitHub project.