Qwen 3.8 27B Uncensored (Thinking)

This abliterated community variant of Qwen 3.8 27B removes safety Refusals from the weights, making it usable for security research and red-teaming — at an MMLU loss of around two points according to the developer. Locally operable under Apache-2.0, with a 262,000-token context and image and video input.

Alibaba Version 3.8 Commercial use permitted Dense 27.8 B 262 K Context 04/2026 locally tested

  • Open Weights
  • Workstation
  • vLLM
  • Text
  • Uncensored
  • Unusable

Sovereign Risk: MEDIUM TODO

LLM Model Review

Updated on · Uncensored

With an overall score of 79.75%, Qwen 3.8 27B Uncensored delivers a remarkably serious performance: a generalist Workstation model with 27.8 billion dense parameters, Thinking mode enabled, and a clear inclination toward analytical work over showmanship. The Speed Profile Badge reads Unusable Tool Expert. That sounds harsher than the overall impression warrants, but it describes the character of this run precisely: strong at reasoning, occasionally solid in tool contexts, yet too sluggish and too unstable in practical response behavior for carefree continuous use. As a multimodal model, fairness requires acknowledging that this text benchmark measures only part of its capabilities — not its image and video side.

Header Metrics: Stability and Reliability

Metric Value Rating Analysis
Timeout Rate 22/49 Unusable The model exhibits catastrophic instability and is entirely unsuitable for unattended production use.
P95 Response Time 389.09 s Critical Extreme tail latency. The model shows massive variance and is unsuitable for time-sensitive processes.

These header metrics are the one finding you cannot write around elegantly. 22 failures in 49 tests are not a cosmetic flaw — they are a real deployment killer. Anyone looking to integrate a local model into agent chains, batch processes, or developer tooling needs one thing above all else: predictability. That is precisely what is missing here. You can forgive good answers. You can even forgive slow answers. But a model that breaks down at a significant rate does not behave like a tool — it behaves like a colleague with no signal.

Architecture and Classification: Much Freedom, Much Responsibility

The pre-assigned category fits the test character surprisingly well. Qwen 3.8 27B Uncensored is classified as a Thinking, Uncensored, Multimodal, Dense, Open Weights, Local, and Tool Use model. This particular run was indeed conducted in Thinking mode, not standard operation. That matters, because longer reasoning paths, more elaborate explanations, and somewhat greater textual weight are not quirks here — they are part of the design.

As a Generalist, the model is not meant to excel at just one thing but to function across the full breadth of tasks. As a Workstation model, robust performance is a reasonable expectation, including in direct comparison with serious mid-tier cloud systems. And because it is dense, the 27.8 billion parameters are not marketing decoration — they are fully active with every response. There is no MoE trick here that dazzles with a massive total count and a small active capacity. That raises the bar. And at its core, the model meets it.

The Uncensored metadata needs to be read carefully. The point is not merely that safety Refusals have been reduced. What matters is whether capabilities collapse in the process. That does not happen here in the benchmark. On the contrary: the model remains remarkably intact in reasoning, code analysis, and structured domain tasks. That is the good news. The bad news is that freedom without discipline means not just fewer restrictions, but also fewer guardrails.

Speed and Token Economy

This local model was evaluated natively on ASUS GX10 / NVIDIA DGX Spark (GB10 Grace Blackwell Superchip, ~115 GB unified memory — no practical memory limit for tested model sizes). The Speed Profile Badge Unusable Tool Expert says plainly: more batch character than conversational flow, more expertise per response than quick interaction, and in this particular run, enough outliers to undermine the appealing Tool Use claim in practice.

This shows up in the flow of responses beyond the bare header metrics. Qwen 3.8 27B Uncensored thinks thoroughly and visibly with purpose. But it does not ramble pointlessly in the process. On the contrary: token efficiency is a genuine strength. No module exceeds the expected verbosity range. CLI, Content, Cultural Intelligence, and UX Writing come in partly well below the fleet median; Documentation lands practically on par, and Code Quality is similarly controlled. The model is not sluggish because it loves the sound of its own voice. The slowness is character, not chatter.

That is a meaningful distinction. A wasteful model costs time because it enjoys talking to itself. This one feels more like a precise but poorly timed engine. The volume of text stays reasonable. The operational stability does not.

Reasoning and Logic: Serious, Correct, Not Quite Elegant

In the Reasoning module, Qwen 3.8 27B Uncensored plays to its Thinking nature convincingly. On the classic two-guards puzzle, it delivers the correct solution, structures the derivation cleanly, stays consistently in German, and presents the logic in a comprehensible form. Particularly effective is the visual presentation with a table and a dedicated reasoning section. This is not a model that merely guesses the right answer. It shows its work.

The weakness lies not in correctness but in the depth of alternative consideration. The Judge plausibly notes that various solution paths are touched on but not fully developed. That is a subtle but important distinction. A good reasoning model solves the problem. A very good one also shows why other paths fail, or which semantically similar questions carry the same logic. Qwen 3.8 27B Uncensored is good here, but not generous.

In the broader picture, the finding remains positive. A Reasoning score of 78.43% is a strong signal for a local Workstation model of this class. The fact that the Uncensored orientation has not dismantled the reasoning capability is not a minor point. Many open derivatives lose their composure precisely there. This one does not.

Code Quality: Technically Sharp, Sometimes Too Broad Rather Than Precise

In the Code Quality area, the model demonstrates why open 27B-class models are now relevant for serious security and review work. The score of 84.52% is no accident. In the security analysis, Qwen 3.8 27B Uncensored reliably identifies core vulnerabilities: SQL Injection, XSS, Session Fixation, Path Traversal, IDOR, CSRF, mail header injection, weak tokens, and type comparison errors. The fix suggestions are technically sound, the German technical terminology is accurate, and the table structure holds up.

What is interesting is the nature of the error. The model does not miss the substance — it miscalibrates. Rather than bundling related vulnerabilities precisely, it splits related issues into too many entries. What should be 19 expected vulnerability groups becomes 24 table rows. That is not wrong in the sense of being nonsensical. It is wrong in the sense of lacking editorial discipline. Security review requires not only seeing a lot, but organizing what you see clearly. Otherwise insight grows into noise.

At the same time, the same test reveals a strength that deserves more weight than a few extra rows. The five implicit or hidden vulnerabilities are not only identified but well explained. The discussion of weak reset tokens, Session Fixation, and secondary injection in particular has real substance. The Judge even commends details where the model pointually surpasses the reference — notably around concrete attack chains involving handler logic that continues running. This is not a bluffer; it is a model with genuine security instinct.

For an Uncensored model, a mitigating context applies: coding and DevSecOps are not its primary purpose. That makes it all the more respectable that in this area it does not present itself as a freewheeling role-player, but as a usable reviewer. The quality reserve is real.

CLI and Tool Use: Strong Planning, a Troubling Tendency to Fabricate

The CLI benchmark comes in at a pleasing 89.0%. That points to solid operational precision on shell and system tasks. Command structures, procedural reasoning, and technical directness appear to suit the model well. For a generalist Workstation model, that is a genuine plus.

In the Tool Use area, the picture is more complicated. The ToolUse score of 78.33% looks robust at first glance. In detail, however, it contains the most dangerous qualitative flaw in the entire report. In one Tool Use task, the model hallucinated content that did not originate from the actual tool result. The P2 score was consequently capped due to hallucination. For content-critical tasks such as research, fact-adjacent reports, or agentic summaries, this is a disqualifying signal.

This is the classic failure mode of modern tool models: they retrieve data, but they do not fully obey it. Rather than deriving conclusions soberly from the tool result, they fill gaps with plausible invention. In small talk, that is annoying. In security, compliance, or research contexts, it is poison. Precisely because Qwen 3.8 27B Uncensored appears competent in so many areas, this misstep is all the more significant. A confidently delivered error is more dangerous than an obviously visible one.

Content Transformation and UX Writing: Between Solid Craft and a Striking Weakness in Language Compliance

The largest qualitative failures lie in Content Transformation. The module score of 71.93% is respectable, but the raw logs reveal why it is not higher. In terms of content, the model is capable. On the video script task, it delivers a professionally structured, engagement-oriented version complete with timestamps, screen annotations, production notes, hook, pattern interrupt, CTA, and even an Easter egg. The material is technically usable.

And then it drives the text into a wall in the wrong language.

In two tasks within the Content module, the model ignored the explicit language instruction and responded in English when German was required. This is not an isolated outlier. Across multiple tasks in the Content area, the model displays a consistent pattern: when faced with simultaneous constraints on language, length, and format, it drops the language requirement first. For production environments with a fixed target language, this is a clear operational risk.

In one of the affected tasks, the model also exceeded the explicit word limit of 900 words, producing 1,524 words169% of the limit. The system applied an automatic deduction of 20%, or 15.72 points, to the achieved partial score. The content quality of the response is therefore irrelevant. The penalty applies regardless. That is exactly how a hard constraint must be treated. A model that cannot follow a brief does not deliver professionally — at best, it delivers something inspired that misses the point.

The language error in those same Content tasks also counts as a Non-Success finding. The model ignored the explicit language instruction and responded in English. This feeds in not merely as a stylistic issue but as a genuine instruction-following weakness. Put differently: the model often knows what to do, but not always which condition is non-negotiable. For marketing and content teams, this is an insidious type of error — impressive in the first minute, embarrassing in the last.

In UX Writing, the picture is milder. At 71.55%, Qwen 3.8 27B Uncensored stays below the top models in the field, but not due to gross errors — rather due to stylistic groundedness. The qualitative sample illustrates this well: from toxic or martial job-listing language, the model produces a functionally clean, safe, and German revision. What is missing is idiomatic sharpness. The reference works with formulations like “Tatkraft und Leidenschaft” or “lösungsorientiert arbeitet sowie Herausforderungen mit Begeisterung begegnet.” The model retreats to more generic terms like “Teamgeist” or “arbeitsfreudige und teamorientierte Persönlichkeit.” That is not bad. It is simply less precisely charged. In editorial terms: no slip-up, but a little too much form and a little too little voice.

Documentation Quality and Cultural Intelligence: Solid, but Not the Most Elegant Pen

At 79.99% in Documentation Quality, the model holds its ground in an area that many open systems needlessly squander. It can structure, explain, and document without burying the audience in jargon. At the same time, it is not a writer with literary ambition. Where the best models sharpen terminology, build didactic arcs, and consciously rhythm their choice of examples, Qwen 3.8 27B Uncensored works more dutifully. The result is usable, rarely brilliant.

In Cultural Intelligence, it reaches 83.6% — clearly above-average composure. The interesting point here is not only the correct defusing of problematic formulations, but the linguistic control. The model keeps German clean, preserves the functional purpose, and remains culturally accessible. At the same time, it stumbles on a detail that is no longer entirely negligible in German-language HR contexts: inclusive formatting such as “in/m/w/d” is missing in one test. That is not a major failure, but it marks the difference between culturally competent and culturally sensitive.

Mode Comparison: Thinking Pays Off Here, but Not for Free

Since standard runs of the same model are also available, Thinking mode can be briefly contextualized. Compared to the standard variants, the overall score of this Thinking run rises visibly to 79.75%, while the standard runs land at 73.86% and 73.38%. That is not a cosmetic difference — it is a genuine performance leap. Particularly notable is that the character becomes clearer as a result: more analytical sharpness, stronger domain performance, better structure.

This gain is not free. The badge remains in the “Unusable” range, and the response character feels more batch-oriented than interactive. That confirms a well-known truth about reasoning models: more thinking often improves the work substantially, but not the pace of the workday. Enabling this mode yields more judgment and costs usability.

Privacy and Data Sovereignty

For this local Open Weights model, there is no cloud dependency in operation. From a European perspective, that is a genuine advantage, since sensitive inputs do not automatically flow to a third-party API provider. The picture around weight provenance is nonetheless not entirely clean: the assessed risk is MEDIUM, because both the base model and the community derivative originate from a Chinese development lineage and no vendor support exists for this variant. For organizations, this means in practice: the runtime path can be local and sovereign, but the supply chain of the weights still demands trust, scrutiny, and ideally internal approval processes.

Conclusion

Qwen 3.8 27B Uncensored is an unusually capable Uncensored local system. It demonstrates that fewer safety Refusals do not automatically mean less competence. In Code Quality, CLI, Reasoning, and Cultural Intelligence, the model shows genuine substance. As a generalist Workstation model with a dense 27.8B architecture, it is not a toy but a serious tool for analytical text work, security review, technical structuring, and more demanding local assistance tasks.

But two hard warning signs must be posted upfront. First: stability in this run is catastrophic. 22 timeouts in 49 tests disqualify the model for unattended agent or production pipelines. Second: Tool Use and content compliance are not clean enough to be trusted blindly. A hallucinated tool derivation and repeated language violations in the Content module are not academic footnotes — they are operational risks.

The right recommendation is therefore not “stay away,” but “only with mature oversight.” Anyone who wants to work locally, prefers Open Weights, and is looking for genuine reasoning-driven quality in Thinking mode will find a remarkably strong package here. Anyone seeking a smooth, reliable workhorse for automated pipelines, however, should not be misled by the score. This model is smart. It is just not yet a clockwork.

This evaluation was generated automatically based on the benchmark data. Model used: GPT-5.4 by OpenAI. The raw data and full methodology are documented in the GitHub project.