Occamy 1.0 35B-A3B (Accio-Lab) (Thinking)

Occamy 1.0 by Accio-Lab is an agentic derivative of the Qwen3.6-35B-A3B checkpoint, focused on long-horizon co-work sessions with tools, structured APIs, and persistent state tracking. The NVFP4 quantization is selective: only the routed experts are quantized, while attention, router, embeddings, and output head remain in BF16. The 35-billion-parameter MoE activates only 3 billion parameters per token and supports 262,000 tokens of context. Apache 2.0 license and documented provenance with recipe, data, and validation artifacts.

Accio-Lab Version 1.0 Commercial use permitted MoE 35 B (3 B active) 262 K Context 12/2025 locally tested

  • Open Weights
  • Workstation
  • vLLM
  • Text
  • Vision
  • Unusable

Sovereign Risk: MEDIUM Occamy 1.0 NVFP4 is a community derivative of Qwen/Qwen3.6-35B-A3B, with a published provenance trail including the quantization recipe, data-provenance file, and validation artifacts. The upstream base is Apache 2.0, the checkpoint runs locally, and the NVFP4 export is limited to routed experts, but Accio-Lab’s organizational jurisdiction is not publicly documented, so the provenance risk remains medium.

LLM Model Review

Created on

With an overall score of 75.25% and the speed profile Unusable Tool Expert, Occamy 1.0 35B-A3B (Accio-Lab) is a model with noticeable ambition — and an equally visible breaking point. As an agentically oriented Workstation model with MoE architecture, 35 billion total parameters, and only 3 billion active parameters per token, it comes across at best as a planning specialist and at worst as a tool operator who can’t quite keep its own tools under control. In the activated Thinking Mode it frequently delivers substance, yet precisely where its profile should shine, competence too often turns into friction. Sovereign Risk: MEDIUM — the weights run locally, but the provenance traces back to a Qwen derivative from within the Alibaba Group ecosystem; the legal origin risk therefore remains a real factor even without cloud telemetry.

Header Notes: Stability and Reliability

Metric Value Rating Analysis
Timeout Rate 20/49 Unusable The model exhibits catastrophic instability and is entirely unsuitable for unsupervised production use.
P95 Response Time 381.01 s Critical Extreme tail latency. The model’s variance is massive, making it unsuitable for time-critical processes.

Architecture and Classification

The pre-assigned categorization hits the mark surprisingly well. Occamy is not a general-purpose chat model but is clearly designed as an agentic reasoning model. Tool use, structured outputs, long working contexts, and persistently conceived sessions are part of the stated specification. The classification as a Workstation model fits accordingly: not a laptop toy, but a serious local system for professional hardware. With this MoE architecture, however, what matters is not the number on the box but the active capacity. Of 35 billion parameters, only 3 billion are active per token. That is the fair benchmark.

This is precisely why expectations need to be calibrated carefully. Occamy presents itself architecturally as large, but internally computes closer to a much smaller active model. This explains why some of its outputs are remarkable while others are surprisingly thin. Anyone reading only the total parameter count expects raw power. Anyone looking at the active parameters understands the character better: specialized, designed for efficiency, but not automatically broadly superior.

The specific test mode also matters. This report explicitly refers to the run with Thinking Mode enabled. For a model targeting multi-step reasoning and agent behavior, this is not a side note — it is the actual stress test. Longer reasoning paths, more verbose responses, and greater friction in terms of speed and stability are therefore not inherently a flaw here. They must, however, pay off in better problem-solving. That is precisely where Occamy becomes both interesting and vulnerable.

Performance in Core Modules

Reasoning and Logic

In reasoning, Occamy shows its most serious side. The logic tasks are not solved elegantly in the sense of maximum didactic clarity, but they are substantively sound. In the metacognition protocol on record, the model correctly works through the classic guard logic, cleanly separates cases, and even supplies alternative framings. This is not a smoke screen of verbose text with little substance, but a genuine reasoning process with a usable result.

The weakness lies more in form than in conclusion. Occamy tends to explain things in prose where a table or a compact proof would be stronger. For human readers this is manageable. For agent pipelines that prefer robust structure and concise verifiability, it is less appealing. The reasoning score is therefore good, but not outstanding: the model thinks correctly, just not always with the precision of a technical writer.

Consistent with the overall picture, Thinking Mode does yield a slight gain over the standard run. The overall score rises from 74.39% to 75.25%, primarily because UX Writing and several reasoning-adjacent areas improve. The price, however, is steep: more output, significantly worse practical stability, and a character that the badge already describes aptly. This does not feel interactive. It feels more like a colleague who fills every whiteboard in the room before answering.

Code Quality and Security

In the Code Quality Audit module, Occamy demonstrates genuine substance. The security analysis of an intentionally vulnerable PHP application is thorough, technically accurate, and sharp in a way one is happy to credit to a model tuned for tool use and orchestration. SQL Injection, XSS, IDOR, CSRF, Session Fixation, weak token generation, Header Injection, hardcoded secrets, Path Traversal: the model identifies not only the obvious holes but also the quieter nastiness in the system. This is not cosmetic checklist work — it is actionable security reading.

Noteworthy is the balance between depth and discipline. The table is cleanly formatted, prioritized by severity, and stays concise within each cell. This is followed by a separate focus on implicit vulnerabilities, including attack scenario and fix. Editorially, this is almost textbook. One minor blemish remains: in the categorization of implicit vulnerabilities, one item is flagged more prominently than it is subsequently elaborated. That is imprecise, but not a security failure.

What is genuinely strong here is the security instinct. Occamy does not hallucinate wildly in the code audit; it names concrete problems with plausible remediation. For security review, code triage, and initial findings preparation, the model is therefore well usable — not as a final authority, but as a technically credible first reviewer. You do not get a penetration testing team in a box. You do get a model that points to the right spots with suspicious frequency.

Content Transformation and UX-Adjacent Writing

In Content Transformation, Occamy performs better than its rather austere name might suggest. The video script protocol shows a model that fulfills its brief cleanly: German language, compact analysis, appropriate spoken-word register, timing markers, production notes, hook, retention elements, CTA, and even a usable Easter egg. This is not merely formally correct — it is genuinely usable as a working draft.

At the same time, the boundary with the top tier is visible. The text is functional but not particularly polished. The visual dramaturgy is simpler than the gold standard, the upfront analysis is deliberately brief, and some of the “why” passages repeat themselves rather than varying the narrative arc. Occamy does not write badly here. It writes like a focused production assistant, not like an author with an instinct for the stage.

In UX Writing and culturally or tonally sensitive rewrites, the model performs solidly to well. The documented HR rewrite case is illustrative: toxic and aggressive phrasing is removed, gender bias is smoothed out, and the tone remains professional. What is missing is a degree of warmth and invitation. Where the better solution says “You are welcome here,” Occamy tends to say “We are looking for.” That is correct, but it breathes less.

Documentation Quality and Cultural Fit

Documentation performance sits in the solid range without producing a standout signature moment in the excerpts reviewed. Combined with the documentation score, this paints the picture of a model that can produce structured technical texts competently, but does not reach the final level of editorial refinement. It is useful, not inspiring. For internal documentation, technical summaries, and system explanations, that is often entirely sufficient.

In the Cultural Intelligence section, Occamy makes few mistakes and gets several things right. The language is appropriate, the regulatory framework is respected, and problematic tonal registers are visibly converted into professional forms. The fact that the resulting text feels somewhat cooler and more corporate than the reference is not a minor point. Especially in HR, brand, or leadership communication, this final layer determines trust. Occamy understands the task. It does not always quite feel it.

CLI and Tool Use

This is where the real fracture in the character profile appears. An agentic model with a tool-use focus may have rough edges in direct tool application — not every shell task needs to look like it came straight from the manual. But Occamy stands out in this area not because of stylistic roughness, but because of a structural problem: the ToolUse sub-score of 42.5 is clearly the model’s Achilles’ heel.

This matters because tool use here is not a side issue — it is part of the architectural narrative. Occamy is designed to plan, structure, and work with APIs and tools. In the benchmark, however, there are two clear warning signs. In at least two tool-use tasks, the originally requested token limit was rejected and the system had to fall back to 4,096 tokens. This is not a trivial detail; it is a header note for any agent deployment. A model that buckles early under larger token requests or contexts loses its most important trust premium precisely in orchestrated workflows.

The more serious finding compounds this: in one tool-use task, Occamy hallucinated content that did not originate from the retrieved tool result. The score was consequently capped by the hallucination cap. For research-oriented, fact-critical, or compliance-adjacent agent work, this is a red warning signal. When a model does not merely interpret tool responses but supplements them as needed, it is no longer an assistant — it is a fabricator with API access. That is precisely what you do not want in this deployment context.

Speed and Efficiency

The local model was evaluated natively on ASUS GX10 / NVIDIA DGX Spark (GB10 Grace Blackwell Superchip, ~115 GB Unified Memory — no practical memory limit for tested model sizes). Its Speed Profile badge reads Unusable Tool Expert. This is not a marketing label; it is almost a summary verdict: for tool-centric, time-critical interaction, generation is too sluggish and the outliers are too large. In practice, this does not mean gentle waiting — it means a workflow that repeatedly runs into a wall.

In sustained operation, Occamy thus presents a contradiction. Token-economically it behaves reasonably overall. No module exceeds the expected verbosity range; several areas even come in below the fleet median. The model does not ramble pointlessly. That is precisely why the weak practical dynamics feel all the more serious. When a model is slow not because of verbosity but remains erratic and sluggish despite disciplined output, that is a runtime characteristic — not merely a prompting issue.

That Thinking Mode costs more than the standard run is expected. But the difference in practical value is limited. The score gain is small, while the operational character becomes noticeably heavier. For users, this simply means: Occamy prefers to think longer rather than deliver faster. That is legitimate — but it then needs to be rewarded with a better answer more often than is the case here.

Hallucinations and Trust Profile

Across large stretches, Occamy is not a habitual fabricator. In security analyses, rewrites, and many transformation tasks, it stays close to the material. That is precisely why the tool-use outlier carries more weight. Hallucinations are not equally damaging everywhere. In a marketing draft, an overreaching detail is annoying. In a response derived from tool results, it is a breach of trust.

The model therefore does not have a pervasive hallucination problem, but it has a critical blind spot at exactly the point where its agentic promise would need to be fulfilled. For fact-sensitive agent work, the rule is therefore: verify, log, double-check. Occamy can work with tools. It must not be trusted blindly — and even less should one blindly trust that it accurately reproduces their results.

Data Privacy and Data Sovereignty

For operational deployment, the situation is double-edged but clearer than with many API models. Occamy is run locally as Open Weights; according to the vendor card, there is no API operation by the developer, and with self-hosting, data remains entirely within the user’s own infrastructure. From a European enterprise perspective, this is a genuine advantage, as no ongoing telemetry to a model provider is described.

Legally, the provenance question does not disappear with that. The publicly documented publisher identity points to a research environment within the Alibaba Group in the People’s Republic of China. The National Intelligence Law, Data Security Law, and Cybersecurity Law are cited as the relevant legal framework of the developer’s parent organization. For European users, this does not mean that locally entered data automatically flows out. It does mean, however, that the provenance of the weights originates from a jurisdiction whose legal situation should not simply be brushed aside in a governance report.

A formally verified GDPR DPA is not documented; data retention is listed as -1 days, meaning it is effectively not applicable to API operation. In practice, this means: self-hosting reduces the data protection risk substantially. Anyone with strict supply chain, provenance, or sovereignty requirements, however, receives a sober and non-trivial warning marker in the calculated Sovereign Risk MEDIUM.

Conclusion

Occamy 1.0 35B-A3B (Accio-Lab) is an interesting, seriously constructed Open Weights model with a clear identity. In Thinking Mode it demonstrates usable reasoning, strong security analyses, solid transformation capability, and overall respectable textual discipline. As an agentic Workstation MoE, it should be measured against its 3 billion active parameters, not the 35B headline. Under that fair benchmark, the performance is genuinely respectable.

The catch, however, is not a small one. Tool use is precisely the area where Occamy stumbles too often: a weak sub-score, token limit fallbacks in tool tasks, and at least one documented hallucination case based on a tool result. Combined with catastrophic overall stability and critical tail latency, this adds up to a model that should not be sent unsupervised into production agent chains. For local security reviews, structured analyses, redrafting, and reasoning-heavy standalone sessions, Occamy can be useful. For autonomous tool-operator roles, it is currently too erratic. The weights provenance is more cleanly documented than many community derivatives, but remains correctly flagged at MEDIUM — not fully cleared — due to the ambiguously documented organizational jurisdiction.

This evaluation was generated automatically based on the benchmark data. Model used: GPT-5.4 by OpenAI. The raw data and full methodology are documented in the GitHub project.