LLM Model Review
Created on
With an overall score of 74.39%, Occamy 1.0 35B-A3B (Accio-Lab) presents a clear profile: not a universal crowd-pleaser, but an agentically oriented Workstation model with Open Weights, 35 billion total parameters, and only 3 billion active parameters per token. That MoE architecture is precisely the benchmark here. Expect not raw Frontier power, but efficient specialization. The evaluated run was conducted in Standard mode, even though the architecture is clearly designed for Reasoning and Thinking. That explains the direct, less expansive tone of many responses. Sovereign Risk: MEDIUM — weight provenance is unusually well documented for a local checkpoint, but Accio-Lab’s organizational jurisdiction has not been cleanly verified publicly.
Header Notes: Stability and Reliability
| Metric | Value | Rating | Analysis |
|---|---|---|---|
| Timeout Rate | 0/49 | Stable | The model ran with absolute stability and reliability throughout testing. |
| P95 Response Time | 66.3 s | Problematic | Significant outliers that interrupt workflow. |
Architecture and Character: What Occamy Wants to Be
Occamy 1.0 35B-A3B (Accio-Lab) is not a standard all-purpose chatbot with a bit of tool glitter on top. The metadata aligns remarkably well with behavior observed in the benchmark. Classified as agentic, the model is visibly trained for multi-step tasks, structured outputs, and tool contexts. As a Workstation model, it is permitted to do more than compact Desktop weights. And as a MoE, what matters is not the large number on the box, but the active capacity of 3 billion parameters per token. That is surgical work rather than brute computational mass.
The consequence is visible in the profile: strong structural fidelity, generally good task decomposition, solid technical matter-of-factness. Less visible are moments of stylistic brilliance or linguistic elegance. Occamy rarely writes poorly, but often writes functionally. It sounds like a system that wants to complete a task, not one that wants to be admired for its prose. For productive tool chains, that is often the right priority.
The second important point is the actual test mode. This report evaluates the Standard run with Thinking disabled. For a model with Thinking and Reasoning DNA, that is no minor detail. In a sense, this tests the street version rather than the race track. Responses come out more compact, visible chains of thought recede, and that is precisely why certain weaknesses must be interpreted with care. The fact that Occamy scores higher in the Thinking run of the same benchmark is no coincidence — it is almost a character confirmation.
Speed: Not a Sprinter, More of a Shift Supervisor
The speed profile reads Interactive DevOps Expert. That sounds like snappy, interactive work in everyday technical contexts, and in practice it only holds with qualifications. Occamy generates at a decent pace on the test system, but not with urgency. Throughput figures are not alarming for a local Workstation MoE, but the long tail does not fit genuine real-time sovereignty. Put differently: for focused sessions the model is serviceable; for hectic ping-pong interaction it lacks the lightness of foot.
This is not entirely atypical for agentic and reasoning-adjacent models. Even in Standard mode, such systems often carry more internal planning overhead than the apparent brevity of their responses suggests. Those who deploy Occamy as a quiet co-pilot for structured DevOps or tool tasks get a workable pace. Those expecting an instantly responsive console will quickly notice that what speaks here is more a deliberate foreman than a jittery terminal acrobat.
Reasoning and Logic: Cleanly Thought Through, Without Theater
In the Reasoning module, Occamy belongs among the pleasantly matter-of-fact systems. The model solves classic logic problems correctly, transparently, and without the false mystique with which some Reasoning models disguise their uncertainty. On the two-guards puzzle it works through the scenarios systematically, correctly explains the double negation, and arrives precisely at the right strategy. Substantively, that is strong. Stylistically, the response could have been more concise and elegant, but here the rule applies: redundantly correct beats brilliantly wrong.
It is notable that the Standard run already exhibits this analytical stability. That confirms that the Thinking/Reasoning classification is not merely marketing wallpaper. Even without Thinking mode enabled, a reasoning-capable core remains visible. Occamy does not argue spectacularly, but reliably. For agentic systems in particular, that is valuable, because planning errors in multi-step workflows are more costly than somewhat longer explanations.
Code Quality and Security: Competent, but Not Forensic
In code- and security-adjacent tasks, Occamy shows perhaps its most practical side. It reliably identifies many classic vulnerabilities, structures findings cleanly in tables, and delivers appropriate fixes. SQL injection, plaintext passwords, XSS, session fixation, CSRF, weak token generation, and insecure cookie boundaries are all correctly addressed. That is no small achievement. Many models deliver attractive headings and empty calories here. Occamy works.
In the security section in particular, however, the limits of its active capacity also become apparent. The Judge records a gap of several unidentified vulnerabilities relative to the reference level. Especially critical is not the overlooking of exotic edge cases, but the underrating of two sensitive issues: Path Traversal and IDOR were classified too low, even though the practical escalation path clearly points toward critical system takeover. That is not a total failure, but it is also not the precision one should accept for serious audits without a second review.
The quality of the fixes is generally serviceable. Prepared statements, password_hash(), password_verify(), hash_equals(), or realpath-based hardening are not window dressing, but functional repair proposals. What is missing is the second layer: exploit chains, attack path synthesis, prioritization by real-world exploitability. Occamy sees many individual holes. It does not always reconstruct the full breach with the necessary sharpness. For Secure Coding Reviews that is good. For offensive or deeply forensic security work, it is not sufficient on its own.
Content Transformation: Strong in Execution, Weaker on Hard Limits
In the Content Transformation module, Occamy delivers results that are in part impressively usable. This is especially true where technical or semi-editorial restructuring is required. A good example is the German video script on two-factor authentication. Here the model works in a structured manner, with functioning timestamps, clear speaker passages, production elements, and an unusually well-developed Easter egg. That is not merely formally complete — it is editorially usable in practice. The response does not sound like a school essay, but like workable production material.
The weaknesses lie less in substance than in discipline. In this module, the model tends to extend a solid framework a step further than necessary, even though the task would already be fulfilled. That is precisely what costs points when hard limits are set.
In one task within the Content Transformation section, the model exceeded the explicit word limit of 250 words by 22%. The system applied an automatic deduction of 20%, or 11.80 points. The substantive quality of the response is therefore irrelevant. The penalty applies regardless.
This finding matters because it fits the model’s agentic character. Occamy tends to solve a task completely and helpfully rather than stopping pedantically at the word boundary. That is almost sympathetic in human terms, but problematic for a machine. In agent frameworks, form pipelines, or publishing systems, a limit is not a friendly suggestion — it is a contract.
Documentation, UX, and Cultural Intelligence: Solid, Rarely Elegant
Documentation quality sits at a good but not outstanding level. Occamy structures cleanly, explains clearly in most cases, and adheres to required formats. It produces no embarrassing walls of text and rarely loses itself in self-repetition. For a local model in particular, that is a strength, because it demands less post-editing in everyday use than many Open Weights of the same class.
In UX writing and microcopy, the model comes across as functional rather than fine-grained. It hits the purpose, but not always the tone. Where concise, pointed, emotionally precise language is called for, Occamy tends to stay in the safe middle lane. That is not wrong. It is simply not the model that suddenly gives an interface personality.
In the area of Cultural Intelligence, the picture is similar. On the positive side, Occamy reliably identifies problematic, exclusionary, or outdated phrasing and converts it into workable German alternatives. In an HR-adjacent rewrite, it removes toxic signals, adds inclusivity, and stays entirely in German. The catch lies in the fine-tuning. It solves the task, but without the cultural or linguistic confidence that turns a correct formulation into a genuinely good modern one. Occamy knows what to avoid. It does not always know how to make the best form out of what remains.
Tool Use and Agentic Suitability: Credible, but Not Flawless
For an agentic model, what matters is not only whether it writes polished responses. More important is whether it handles structured outputs, tool proximity, and longer workflows. In this regard, Occamy delivers a credible profile. The Model Card promises a focus on tool use, APIs, state tracking, and longer co-work sessions. The benchmark does not contradict this. One sees a system that organizes tasks, takes structured formats seriously, and thinks in technically purposeful terms.
At the same time, the ToolUse score is no cause for celebration. This suggests that while Occamy smells like an agent, it does not always achieve the precision of a specialized tool executor in direct tool execution. That must be read fairly: agentic models are often built to plan tasks and shape subtasks, not to execute every exact individual step perfectly as a one-shot. Nevertheless, it remains the case that the bar for a tool-use model is high. Those who integrate it into real agent chains should carefully guard calls, structures, and abort paths.
Token Economy: Pleasantly Mature
Occamy behaves token-economically — no module exceeds the expected verbosity envelope. For a local model, that is more than a stylistic question. Fewer output tokens generally mean shorter wait times and less friction in the workflow. Particularly positive is the fact that even more extensive tasks in Code Quality, Content Transformation, and UX do not tip over into pointless embellishment.
This efficiency makes the previously mentioned word-limit violation all the more interesting. It was not a symptom of general verbosity, but a localized loss of control at a clear hard boundary. That is almost the better news, because it seems more manageable at the prompt and control level than fundamentally vague model behavior.
Data Privacy and Data Sovereignty
Not applicable as a separate risk block, since Occamy 1.0 35B-A3B (Accio-Lab) is operated here as a local Open Weights model. What is relevant instead is the provenance of the weights: it is unusually well documented for a community derivative, but remains a case for informed caution rather than blind trust, given Accio-Lab’s organizational jurisdiction not being clearly verifiable publicly.
Conclusion
Occamy 1.0 35B-A3B (Accio-Lab) is a local model, evaluated natively on the ASUS GX10 / NVIDIA DGX Spark (GB10 Grace Blackwell Superchip, ~115 GB Unified Memory — no practical memory limit for tested model sizes), and it makes sense in precisely that context: as a serious agentic Workstation candidate with Open Weights, clean MoE efficiency, and sound technical discipline. It is strong on logic, structured analysis, security basics, and production-adjacent content transformation. It weakens where absolute precision in severity ratings, strict constraint adherence, or stylistic finesse is required. The Standard run reaches 74.39%; the Thinking run of the same model scores visibly higher and fits the model’s own architecture more closely. Those using Occamy should therefore be aware that what was tested here is the throttled everyday mode, not the fully deployed Reasoning variant. Across all tests, no notable hallucinations — the model prefers to invent little rather than embarrass itself with grand gestures.
The recommendation is therefore differentiated, but clear. For local assistance in DevOps-adjacent workflows, structured content work, technical editorial tasks, and as a component in agentic chains, Occamy is a serious option. For high-stakes security audits, strictly formatted publishing pipelines without post-review, or tasks where every word limit carries legal weight, tighter guardrails are needed. Occamy is not a bluffer. It is a willing worker with substance. But give it the workbench — not the corner office.
This evaluation was generated automatically based on the benchmark data. Model used: GPT-5.4 by OpenAI. The raw data and full methodology are documented in the GitHub project.