Signal 3.8 27B

Signal 3.8 27B by AgentionAI is a token-efficiency fine-tune on Qwen 3.8 27B: via self-distillation, the model generates roughly 57 percent fewer response tokens and thinking chains half the length, at equal or better quality. Licensed under Apache 2.0, available in a mid-range Q5 quantization with a 262,000-token context and local inference.

AgentionAI Version 3.8 Commercial use permitted Dense 27.8 B (27.8 B active) 262 K Context 12/2025 locally tested

  • Open Weights
  • Workstation
  • llama.cpp
  • Text
  • Vision
  • Instruction-Tuned
  • Batch

Sovereign Risk: MEDIUM TODO

LLM Model Review

Created on · Instruction-Tuned

With an overall score of 73.23%, Signal 3.8 27B presents a clearly readable profile: a tightly trained generalist in the Workstation class with 27.8 billion dense parameters, one that prefers to answer directly rather than get lost in its own text. The Speed Profile Badge “Batch DevOps Expert” fits surprisingly well: this model feels less like a nervous chat sprinter and more like a methodical batch worker for structured tasks, code audits, and tool-adjacent workflows. The operating mode matters: what was tested here is Standard — i.e., with Thinking disabled — even though the architecture supports Extended Thinking in principle. Sovereign Risk: MEDIUM — the weights originate from a Chinese development environment or a Chinese fine-tune context; since deployment is local, the primary risk here is less about data exfiltration than about provenance and supply chain.

Header Notes: Stability and Reliability

Metric Value Rating Analysis
Timeout Rate 7/49 Unreliable The model is unreliable and drops out significantly often in practice.
P95 Response Time 158.13 s Critical Extreme tail latency. The model exhibits massive variance and is unsuitable for time-critical processes.

The header notes are the part you cannot argue away in day-to-day use. They damage the practical impression more than any elegant individual score. Because Signal 3.8 27B is often better on substance than these stability figures suggest. But a local model that exhibits massive variance in the long tail of responses and repeatedly drops out is only suitable for productive pipelines if retries, watchdogs, and sensible timeout strategies are already part of the operational concept. For interactive one-on-one conversations, that is a nuisance. For agent pipelines, it is a genuine operational concern.

Architecture and Character: A Tightly Tuned Instruct Generalist with Reserves

The metadata here is not a side note — it is the key to proper classification. Signal 3.8 27B is recognizably trained as an Instruct model for direct command execution. That is exactly what the tests show. Responses tend to stay focused, format-aligned, and relatively concise. The fact that the architecture is Thinking-Optional should not be confused with this particular test run: the model could reason more extensively, but did not do so here because the run was explicitly conducted in Standard mode. That is not a deficiency — it is the chosen operating state.

As a Generalist, it must account for itself across the full breadth of tasks. As a Workstation-class model, more can be demanded of it than of a 7B or 12B candidate. And as a Dense architecture, every one of the 27.8 billion parameters counts toward active capacity. There is no MoE excuse here, no hidden shrinkage. What you get is raw, always-active model size. That is precisely why weaknesses in reasoning polish, code completeness, and reliability stand out more visibly. This model is not undersized. It is deliberately conditioned.

Speed: Not a Dialogue Foil, More of a Serial Worker

Signal 3.8 27B was evaluated as a local model on an NVIDIA DGX Spark (GB10 Grace Blackwell Superchip, ~115 GB Unified Memory — no practical memory limit for tested model sizes). Its badge “Batch DevOps Expert” describes the character aptly: moderate to low response throughput, but a profile that suits stackable, structured tasks better than ultra-direct chat interaction.

In running text, one thing stands out above all: the model is token-economical. No module exceeds the expected verbosity range. On the contrary, in CLI, Code Quality, UX Writing, Cultural Intelligence, and Documentation, Signal often stays well below the fleet median. That is the good news. The bad news is that token frugality does not compensate for weak tail stability. Writing shorter helps, but it does not fix outliers. Signal is therefore not slow because it rambles. It is slow and high-variance despite being comparatively concise. That is an important distinction.

Code Quality and Security: Sharp in Findings, Too Often Incomplete as a Package

The strongest functional area is, of all things, the most sober one. In Code Quality and Security Audit, Signal 3.8 27B works with solid technical substance. In the security review at hand, it identifies 19 vulnerabilities, reliably hits central findings such as SQL Injection, plaintext passwords, Path Traversal, IDOR, and loose API key validation, and delivers the required Markdown table in formally correct shape. That is no small feat. Many models fail here either on completeness or on format discipline.

What is interesting is the character of the failure. Signal does not fail as a bluffer — it fails as an abbreviator. The Judge sees the core analysis as “80 to 85 percent of the Golden Standard.” What is missing is not the obvious gaps, but two concrete standalone entries in the table, including an explicit CSRF vulnerability and a reset token without expiry as a separate finding. Also absent is a narrative attack chain and a concluding assessment of production readiness. In other words: the model recognizes a great deal, but packages its findings too tersely. For a security engineer, that is correctable. For an inexperienced user, this very terseness can be dangerous, because it makes the difference between “good analysis” and “complete analysis” invisible.

In a security context, that matters. A model that cleanly surfaces five implicit or advanced vulnerabilities deserves respect. A model that then delivers no complete prioritization chain and no closing verdict still falls short of its potential. Signal is no botcher here. It is more like the auditor who steps away from the whiteboard too soon.

CLI and Tool-Use: Surprisingly Robust Where Others Stumble

The Tool-Use score and the strong CLI findings speak to a model that is more useful in agentic or semi-agentic workflows than the overall impression initially suggests. This also aligns with the tag combination of Tool-Use plus Instruct. Such models do not need to formulate brilliantly — they need to execute precisely. That is exactly where Signal succeeds in the benchmark’s system-proximity tasks, often more so than in the representative writing disciplines.

The “Batch DevOps Expert” badge gains substance here: Signal behaves like a model that handles shell-adjacent, structured, technical work orders competently, as long as you do not simultaneously demand rhetorical elegance, maximum completeness, and perfect user guidance. That is not a romantic form of intelligence. It is the useful kind.

Reasoning and Logic: Correct, but Without Distinction

Reasoning reveals the central tension of this model. The architecture is fundamentally capable of multi-step thinking. However, testing was conducted without Thinking activated. Accordingly, no elaborate chains of thought should be expected. Even so, the result is only half satisfying. On the positive side: in the guardian task, Signal arrives at the correct solution — the logic holds, the answer is complete. On the negative side: the presentation is redundant, less clearly structured than necessary, and harder to follow in its explanation than stronger models.

The verdict therefore requires careful calibration. This reasoning is not weak in the strict sense. It is correct, but not particularly elegant. The model apparently thinks enough to avoid failure, but not visibly enough or not structuredly enough to give the reader the sense that a particularly sharp logical instrument is at work. For standard operation, that is acceptable. Anyone selecting Signal for its optional Thinking capability will only discover in activated Thinking mode whether more architecture actually delivers more clarity.

UX Writing, Content Transformation, and Documentation: Functional, but Rarely Inspiring

In the text-adjacent modules, the training philosophy becomes visible. Signal is conditioned for concise, direct responses. That helps with tables, audits, and clear work orders. It hurts where formulation is not merely a carrier but part of the task itself.

Cultural Intelligence provides a good example. The rewrite of a toxic job posting succeeds functionally: in German, professional, free of toxic language, without reproducing bias, without meta-commentary. That is the baseline. In the higher register, however, energy is missing. The Judge describes the text in effect as a correct HR document rather than a genuinely inviting job posting. The model replaces harshness with bureaucracy. It defuses, but does not inspire. Those seeking employer branding get administrative language with manners.

Similarly in Documentation Quality and UX Writing: the scores are not catastrophic, but they do not reveal a model that understands form as a competency in its own right. Signal writes readably, often tidily, sometimes even pleasantly concisely. It rarely writes with that combination of precision and warmth that distinguishes excellent product copy. That is a difference you only notice when users drop off or teams start routinely rewriting the outputs.

In the Content Transformation module, Signal delivers a neat paradox. Substantively, the result is strong. The video script adaptation for a 2FA tutorial is rated by the Judge as production-ready, with usable timecodes, clear visual cues, spoken-word-appropriate style, and cleanly integrated production cues. Even the Easter Egg is implemented creatively. Yet the model takes a hard hit here — not for substantive weakness, but for a lack of discipline around the word limit.

In a task within the Content Transformation section, the model exceeded the explicit word limit of 900 words by 21%. The system applied an automatic deduction of 20%, or 16.12 points. The substantive quality of the response is therefore irrelevant. The penalty applies regardless. That is precisely the point: Signal can deliver good work and still knock itself off course with a formal rule violation. Anyone working with fixed editorial templates, character limits, or UI constraints should not dismiss this as a trivial error. In such environments, form is not decoration. It is a contract.

Cultural Intelligence: Polite, Inclusive, but Somewhat Bloodless

The Cultural Intelligence score ranks among the stronger areas. The model works cleanly in German, reliably removes problematic phrasing, and shows no gross missteps. Particularly important: it adheres to the instruction without drifting into moralizing explanatory text. That is by no means a given when rewriting sensitive content.

And yet a faint aftertaste remains. The qualitative comparison shows that Signal conveys positive values such as initiative, enthusiasm, and an inviting tone less distinctly than the reference. It can defuse, but is less adept at reframing. It removes the poison from the text without automatically making it come alive. Strong for compliance. Merely adequate for communication.

Hallucinations: Pleasingly Low Nonsense Factor

One advantage runs across all modules: Signal 3.8 27B does not behave like a model that would rather invent something when in doubt. Its weaknesses lie in conciseness, structure, and formal discipline — not in fabricated facts or freely assembled security findings. Particularly in Security, Documentation, and transformation-heavy tasks, that counts for a great deal. A sober model with gaps is easier to manage than an eloquent one that sells nonsense with style.

Conclusion

Signal 3.8 27B is an interesting case of deliberately trained sobriety. As a Generalist in the Workstation class with a Dense architecture, it delivers in Standard mode a profile that convinces primarily in code audit, CLI proximity, and structured task execution, while showing visible weaknesses in UX-adjacent language, motivating tone, and formal length precision. The token-economical tuning is no marketing fiction — it is clearly recognizable in the benchmark. The model writes concisely, often pleasantly so. But brevity is no longer a virtue when it comes at the cost of completeness, warmth, or rule adherence.

For productive deployment, the real question is not quality but reliability. The substantive performance is sufficient for serious work in several modules. However, the instability and critical tail latency draw a hard line: unsupervised in time-critical pipelines, Signal is not a good idea in this run. With retries, clear prompt framing, and human review, it can be a serviceable local tool — especially for technical teams that value concise, direct output.

Local deployment also brings the sovereignty dimension: Apache 2.0 license, local weights, no cloud dependency. That is a genuine advantage. The weights provenance risk remains medium, however, because both the base model and the fine-tune originate from a Chinese development environment and the training disclosures have not been independently verified. Across all tests, no noteworthy hallucinations. The model would rather produce nothing than embarrass itself. Compared to the separately available Thinking run of the same family, this Standard run shows the expected character difference: more direct, more concise, less reasoning-driven, and weaker overall. Anyone selecting Signal should therefore be very deliberate about distinguishing between Standard for lean execution and Thinking for more demanding logic. In its current form, it is no charismatic performer. But it is a serviceable, technically serious worker with poor manners when it comes to waiting.

This evaluation was generated automatically based on the benchmark data. Model used: GPT-5.4 by OpenAI. The raw data and full methodology are documented in the GitHub project.