Qwen 3.8 27B Uncensored

This abliterated community variant of Qwen 3.8 27B removes safety Refusals from the weights, making it usable for security research and red-teaming — at an MMLU loss of around two points according to the developer. Locally operable under Apache-2.0, with a 262,000-token context and image and video input.

Alibaba Version 3.8 Commercial use permitted Dense 27.8 B 262 K Context 04/2026 locally tested

  • Open Weights
  • Workstation
  • vLLM
  • Text
  • Uncensored
  • Batch

Sovereign Risk: MEDIUM TODO

LLM Model Review

Created on · Uncensored

With an overall score of 73.86%, Qwen 3.8 27B Uncensored presents a profile that cannot be dismissed with a shrug: a local generalist in the Workstation class with 27.8 billion dense parameters, a multimodal foundation, and tool ambitions — tested here, however, in standard mode without Thinking enabled. Its Speed Profile Badge is Batch Tool Expert. That fits the character of this model remarkably well: more methodical worker than quick-fire sparring partner, with clear strengths in structure, security, and reasoning, but a troubling tendency to switch languages under instruction pressure or to hallucinate facts in tool output. Sovereign Risk: HIGH — the provided cards place the vendor under Chinese law with processing in China; for European organizations, that is a real sovereignty and compliance factor.

Header Scores: Stability and Reliability

Metric Value Rating Analysis
Timeout Rate 10/49 Unreliable The model is unreliable and drops out significantly often in practice.
P95 Response Time 168.06 s Critical Extreme tail latency. The model scatters massively and is unsuitable for time-critical processes.

These header scores set the tone right away. Ten timeouts in 49 tests are not background noise — they are a practical problem. Anyone planning to deploy a local model in agent workflows, batch pipelines, or semi-automated editorial pipelines needs to budget for retries, watchdogs, and patience. Add to that a critical tail latency. In other words: even when the model responds well, in a meaningful share of cases you wait too long for it. For a model with batch character, that may still be acceptable in some documentation or analysis workloads. For interactive use, it is unpleasant.

Architecture and Classification

The pre-assigned category hits the mark with surprising precision. Qwen 3.8 27B Uncensored is classified as a Thinking architecture, but was explicitly tested here in standard mode. That matters, because the more concise, direct answers from this run should not be read as a deficiency. The capacity for deeper inference is present — just not unlocked. At the same time, it is a Generalist, not a pure code model and not a specialized vision-language system, even though the multimodal foundation limits what the benchmark can say: the text benchmark at hand measures only part of its actual breadth.

As a Dense model, it must be held to its full 27.8 billion parameters. There is no MoE discount here where only a small expert slice would be active. For the Workstation class, one is therefore entitled to expect a robust performance level across multiple disciplines. Qwen delivers on that in part. What is disruptive is not a lack of basic capability, but breakdowns in discipline and reliability. That is particularly relevant for an uncensored derivative: freedom without internal order is not a feature — it is often just another name for risk.

Speed and Token Discipline

On the local reference system ASUS GX10 / NVIDIA DGX Spark (GB10 Grace Blackwell Superchip, ~115 GB Unified Memory — no practical memory limit for tested model sizes), Qwen 3.8 27B Uncensored presents itself as a clear Batch Tool Expert. This badge denotes a model that does not aim for snappy real-time response, but rather for longer, task-oriented runs with a tool focus. In practice, that means: suited for tasks where you kick off a run and collect the result later. For tight-cadence dialogue, this is the wrong temperament class.

The model does work in a token-economical manner. No module exceeds the expected verbosity envelope. For a local model, that is more than a cosmetic point, because concise output directly reduces perceived response time. The discipline is particularly noticeable in CLI, Cultural Intelligence, and UX Writing. Qwen does not beat around the bush unnecessarily there. When it stumbles, it is usually not from verbosity but from miscalibration.

Code Quality and Security: Substance Over Show

The strongest side of this model lies in technical auditing. In the code quality area, Qwen 3.8 27B Uncensored delivers a security analysis worth taking seriously: cleanly structured Markdown table, correct categorization by severity and type, viable fixes ranging from prepared statements to session hardening. Particularly positive is the detection of implicit vulnerabilities as well, including Mail Header Injection, Type Juggling, IDOR, and Second-Order SQL Injection. This is not decoration for bug bounty slides — it is practical substance.

It is not entirely flawless. According to the Judge, some entries are missing — for example, separately listed hardcoded credentials, an API secret, or individual vulnerability lines that the gold standard explicitly enumerates. That affects completeness, not technical seriousness. The key point: the fixes identified are largely correct and directly actionable. The model knows its material.

For an uncensored derivative, that is noteworthy. Such variants tend to lose internal consistency when safety vectors are surgically altered. In the coding core, remarkably little of that is visible here. Qwen does not appear damaged in the security audit — it appears largely intact. That deserves respect.

Reasoning and Logic: Sharp, but Not Elegant

In the logic and reasoning area, Qwen 3.8 27B Uncensored demonstrates why the Thinking classification is not pulled from thin air. Even in the standard run, the model argues correctly, breaks down case distinctions cleanly, and arrives at the right solution on classic logic tasks. The judge logs describe the reasoning as substantively correct, detailed, and coherent.

The catch is in the form. Qwen solves the problem, but not with textbook precision. The presentation is repetitive, partly self-correcting, sometimes a bit roundabout. You can almost sense that more reasoning capacity is present than the chosen mode draws on. The result is right; the path to it is just not always well-paved. Those seeking logical correctness will find it. Those expecting didactic elegance will need to make concessions.

Especially in comparison with the Thinking run of the same model family, the difference in character becomes apparent. The Thinking variant achieves visibly greater overall pressure and fits the reasoning architecture better. The standard run, by contrast, feels more compact but also somewhat held back. That is not a measurement error — it is precisely the point of a dual run.

UX Writing: Solid User Guidance Without Academic Scaffolding

In UX Writing, Qwen delivers pleasingly solid craft. The optimization suggestions are in German, user-centered, and organized in a usable table structure. The model hits tone, format, and practical applicability cleanly. The judges praise the recommendations as professional and directly deployable. That is exactly how a generalist should perform in this module.

What is missing is the final layer of theoretical grounding and validation strategy. Qwen writes like a good product person, not like a senior researcher with a literature apparatus. That is no flaw for day-to-day work, but it is the difference between “good” and “excellent.” The text is useful. It simply makes no claim to methodological rigor.

Content Transformation: Strong Craft, Weak Discipline

This is where the model becomes simultaneously interesting and frustrating. In a demanding video script task, Qwen 3.8 27B Uncensored produces a substantively strong result: timestamps, screen annotations, production cues, retention hooks, an Easter egg, a CTA. The Judge explicitly describes the script as production-ready. That is the good news.

The bad news is more serious. In this task, the model ignored the explicit language instruction and wrote the actual script portion predominantly in English, despite German being required. It also exceeded the clear word limit of 900 words, reaching 1,127 words125% of the limit. The system applied an automatic deduction of 20%, or 14.84 points, against the achievable score. The substantive quality of the response is therefore irrelevant. The penalty applies regardless.

The model ignored the explicit language instruction and responded in English in a task in the Content Transformation area. In production environments with a fixed target language, that is a clear deployment risk.

This finding is not merely an embarrassing slip. It reveals a structural weakness: under simultaneous constraints of language, length, and format, Qwen loses instruction discipline first. Content Transformation is precisely the module where a model must juggle multiple requirements at once. When it drops the target language and the word limit in exactly that context, it is not a cosmetic flaw. It is a workflow risk.

Documentation Quality: Substantively Usable, Linguistically Not Always Compliant

The documentation area shows the same split picture. On one side, Qwen can deliver comprehensive, structured responses at the right level of detail. Token usage is disciplined, which helps with long documentation tasks. On the other side, a language error recurs: in a documentation task, the model responded in English despite an explicit German instruction.

The model ignored the explicit language instruction and responded in English in a task in the Documentation Quality area. This is not a technical defect — it is a weakness in instruction following.

Because the same error type appears in two different modules, the language failure is not an isolated outlier. Across multiple tasks in the content and documentation areas, the model shows a consistent pattern: under simultaneous constraints of language, length, and format, it drops the language requirement first. For German-speaking organizations or editorial teams, this is not a theoretical flaw — it is an immediate acceptance risk.

Cultural Intelligence: Sensitive Enough, Just a Bit Bureaucratic

In the Cultural Intelligence module, Qwen performs adequately to well. A toxic job posting is correctly detoxified, rewritten in gender-inclusive language, and fully rendered in German, without explanatory ballast or format violations. The model understands how to defuse culturally sensitive language without mangling the meaning.

The qualitative gap from the top group lies less in compliance than in tone calibration. Where stronger models build in warmth, motivation, and idiomatic elegance, Qwen sounds somewhat more sober — almost administrative in register. That is functionally perfectly fine. It is just not the kind of text that makes you spontaneously think: this is how I want to work. Good HR department, not much poetic pulse.

CLI and Tool Use: Solid Execution, Questionable Trust Basis

In the CLI area, Qwen stands on solid ground. That fits the badge: tool-oriented tasks are fundamentally in this model’s wheelhouse. But this is also where one of the most sensitive findings of the entire benchmark sits. In a tool use task, the model produced content that did not originate from the retrieved tool result but was hallucinated. The score was capped by a hallucination limit.

For content-critical tasks, that is a disqualifying signal. A model using a tool may condense, organize, and explain its result. It may not supplement it as though reality had delivered a bonus track. In research, forensics, incident response, or reporting, exactly this behavior is dangerous. A tool use model that occasionally treats its tools as suggestions is like an accountant with poetic license.

Privacy and Data Sovereignty

The card situation is contradictory, but not without consequence. The model tested here runs locally and under Apache 2.0, which is a strong argument for data sovereignty in practical deployment. At the same time, the model card attributes the base to Alibaba, and the provided vendor card references Alibaba Cloud Computing Ltd., headquartered in Hangzhou, China, with applicable law China (PIPL/CSL/DSL) and data location China. The calculated sovereignty value is HIGH, citing Chinese legal access rights. For EU users, the sober reading is: as soon as this model were used not locally but via the named vendor infrastructure, an elevated transfer and access risk would arise. A verified GDPR DPA status is not in place; the data retention period is specified as -1 days, which is not meaningfully defined. For regulated European organizations, this constitutes a compliance obstacle in cloud form. For the local deployment tested here, the most relevant factor is primarily the MEDIUM Weights Provenance Risk: open weights, yes — but with a supply chain one should know and consciously accept.

Conclusion

Qwen 3.8 27B Uncensored is a contradictorily capable model. As a local, open Workstation-class generalist, it brings genuine technical competence — particularly in security audits, structured reasoning, and formally usable work outputs. The multimodal and tool-capable foundation extends the theoretical deployment space, even if this benchmark illuminates only the text portion. The Apache 2.0 license and local deployability make it attractive for sovereignty-oriented teams, despite a medium provenance risk on the weights.

But one should not be lulled by the raw score. The header scores are rough. Ten timeouts, critical tail latency, repeated language violations, and a documented tool hallucination paint the picture of a model that requires supervision. For security reviews, technical analyses, red-teaming, and longer batch tasks, it is a sensible choice. For language-critical publishing workflows, fact-bound tool pipelines, or unsupervised agent runs, caution is mandatory. In direct family comparison, the standard run feels more sober and less fully realized than the Thinking variant of the same model, which delivers clearly greater overall performance pressure. The character remains the same throughout: capable, willful, not well-behaved. That can be an advantage. Without guardrails, it is often simply exhausting.

This evaluation was generated automatically based on the benchmark data. Model used: GPT-5.4 by OpenAI. The raw data and full methodology are documented in the GitHub project.