Command A+

Cohere’s first open MoE model combines frontier-level performance with self-hosting: of 218 billion total parameters, only 25 billion are active per token, the context window spans 128,000 tokens, and the model processes both text and image inputs. Released under the Apache 2.0 license out of Canada, with no vendor lock-in and no restricted weights.

Cohere Version 05-2026 Commercial use permitted MoE 218 B (25 B active) 128 K Context 12/2025

  • Open Weights
  • Frontier
  • Cohere
  • Text
  • Vision
  • Multilingual
  • Long Context
  • Real-Time

Sovereign Risk: LOW Cohere is a Canadian provider (Toronto). Command A+ is licensed under Apache 2.0 and available as an Open Weights model on Hugging Face. No US jurisdiction, no proprietary Restricted Weights risk. Weights are freely downloadable and self-hostable.

Tool-use profile: 6 assets in detail

Comparison of asset performance (P1/P2/Combined) against the fleet average

Asset performance (radar)

Score Breakdown vs. Fleet Average


Tool-use details

Asset performance, reliability, and runtime profile

CrucibleMark evaluates tool use across 6 independent tests. Click a test name for details.

Reliability

  • Tool Call Valid: No
  • Retry: Not required
  • Hallucination: Not detected

Reliability measures how consistently a model actually executes tool calls: Tool Call Valid schema and format accepted, Retry Required successful only after retry, Hallucination Flag fabricated tools or parameters detected. All three green means production-ready.

Runtime profile

Token
0
Input + Output
Cost
$0
Cost per Run

The runtime profile shows the latency and cost metrics for the model run: Call 1 First Request, MCP Protocol Latency, Synthesis Response Generation, Total sum of all phases, supplemented by Token input and output and Cost cost per run.

Tool-use review

Created · Multilingual · Long Context

Deployment Verdict

Created on: 22.06.2026, 21:36:17

Conditional deploy, because no usable execution data exists for the production-critical signals on tool usage, and the combined score of 0.00 does not permit an evidence-based release. The only positive is that no hallucination event was detected. That does not substitute for a passed tool run.

Tool Execution Profile

There is no reliable data for Command A+ on whether it selects the right tool, produces valid calls, or operates in an MCP-compliant manner. That is the core finding here. For a model positioned as an agentic Frontier system, one would expect the Web Search and Tool Selection test to reveal whether it situationally distinguishes between search and direct fetch, and the URL Construction test to show whether it derives target addresses precisely enough for deterministic workflows. Both are missing. It is therefore impossible to assess whether the model solves tool selection as a planning problem or merely reproduces a rigid call pattern. Since no retry was required, there is also no indication of a pure formatting issue. Execution evidence is simply absent.

Synthesis Fidelity

How well does it condense tool results? No P2 data exists for this. For production decisions, that is a hard gap — because the actual value creation in an MCP pipeline lies not in the tool call itself, but in the reliable condensation of its returns. A strong model must compress source content without losing structure, constraints, or boundary conditions. This remains open.

Does it stay within the tool result or fall back on training data? No data is available for EU License Research, the honeypot test for current license lookups versus training knowledge. At least no hallucination flag was set. That is a weakly positive signal, but not a proof of trustworthiness. Without a honeypot result, it remains untested whether the model stays cleanly bound to tool output in compliance-adjacent pipelines.

Error Resilience

No data exists for the 404 test on how the model responds to failing tool calls. It is therefore unclear whether Command A+ reports errors transparently, asks clarifying questions, or fabricates substitute content despite the failure. That dividing line is precisely what matters in production. A model may be incomplete when a tool fails. It must not speculate.

Operational Profile

No reliable latency or per-run cost data available. Local operation is possible. Cost-effectiveness relative to performance remains open without measured values.

Conclusion & Recommendation

Command A+ remains for now a candidate with solid deployment properties on paper: openly licensed, locally operable, long context window, agentic orientation. For a real MCP tool pipeline, that is not enough. I would only put it into a controlled pilot with tight observability, enforced tool call validation, and clear fallbacks. Not recommended for release in compliance-, research-, or fetch-heavy production paths without manual oversight at this time.

This analysis was generated automatically based on benchmark data. Model used: GPT 4.5 by OpenAI. The raw data and full methodology are documented in the GitHub project.