Kimi K3

Kimi K3 has been Moonshot AI’s Frontier flagship since July 2026, with 2.8 trillion total parameters making it the largest Open Weights model in the world to date. The Stable-LatentMoE architecture activates only 16 billion parameters per token from 896 experts; the context window spans 1,000,000 tokens for long-horizon agentic workflows. Kimi Delta Attention and Attention Residuals are designed to significantly improve scaling efficiency over the predecessor model.

Moonshot AI Version k3 Commercial use permitted MoE 2800 B (16 B active) 1000 K Context 02/2026 $3 / $15 per 1M

  • Open Weights
  • Frontier
  • OpenRouter
  • Text
  • Vision
  • Video
  • Long Context
  • Agentic Orchestrator
  • Batch

Sovereign Risk: HIGH Moonshot AI is a Chinese company subject to China’s National Security Law (NSL), which may enable state access to data via services hosted in China. In February 2025, Germany’s BSI explicitly warned against the use of Chinese AI cloud services; this risk assessment applies here conservatively as well. Kimi K3 has been announced as a 2.8T Open Weights model whose weights are scheduled for release on July 27, 2026; the operational risk arises primarily from cloud usage under Chinese jurisdiction, not from the mere existence of open weights.[web:605][web:610]

Tool-use profile: 6 assets in detail

Comparison of asset performance (P1/P2/Combined) against the fleet average

Asset performance (radar)

Score Breakdown vs. Fleet Average


Tool-use details

Asset performance, reliability, and runtime profile

CrucibleMark evaluates tool use across 6 independent tests. Click a test name for details.

Reliability

  • Tool Call Valid: No
  • Retry: Not required
  • Hallucination: Not detected

Reliability measures how consistently a model actually executes tool calls: Tool Call Valid schema and format accepted, Retry Required successful only after retry, Hallucination Flag fabricated tools or parameters detected. All three green means production-ready.

Runtime profile

Call 1
5.26
First Request
MCP
0.92
Protocol Latency
Synthesis
50
Response Generation
Total
337.15
Sum of All Phases
Token
19349
Input + Output
Cost
$0
Cost per Run

The runtime profile shows the latency and cost metrics for the model run: Call 1 First Request, MCP Protocol Latency, Synthesis Response Generation, Total sum of all phases, supplemented by Token input and output and Cost cost per run.

Tool-use review

Updated · Long Context · Agentic Orchestrator

Deployment Verdict

Conditional deploy, because Kimi K3 is strong at tool selection and showed no hallucination during the run, but tool calls were not consistently valid and synthesis quality remains only moderately stable for reliable production handoffs.

Tool Execution Profile

Kimi K3 demonstrates genuine tool intelligence rather than mere template usage. In the Web Search & Tool Selection test — which checks whether web_search is chosen over fetch without any hint — it makes the right call cleanly. This speaks to workable orchestration in open pipelines. In the URL Construction test, which measures autonomous derivation of a target URL followed by a fetch, it remains usable but not deterministic enough for systems that depend on precise call formats. The overall finding is therefore split: good planning logic, but no consistently clean MCP execution. The fact that the tool call was flagged as invalid is relevant for productive tool chains. It points less to a lack of task understanding than to execution discipline at the protocol level.

Synthesis Fidelity

How well does it consolidate tool results? Solid, but not strong enough for high-quality decision summaries. The broader research tasks fall off noticeably: EU License Research and Multilingual Search & Synthesis land at only 60 in consolidation. Kimi K3 can merge retrieved content but loses precision and prioritization in the process. For operational responses this is often still acceptable. For compliance, policy, or executive summaries it is too imprecise.

Does it stay within tool results or fall back on training? Here the model is more trustworthy than the P2 score might suggest. In the honeypot EU License Research — which checks whether current license restrictions are answered from web sources rather than training knowledge — no hallucination was detected. This is the central trust signal of this run.

Error Resilience

Acceptable for production. In the Tool Failure Handling (404) test, which checks for transparent behavior on a failing tool call versus fabricated fallback content, Kimi K3 communicates the failure without inventing page content. That is exactly the minimum requirement for tool pipelines. A model is allowed to fail. It must not conceal the failure.

Operational Profile

Total 337.15s per run. Call 1: 5.26s. MCP latency: 0.92s. Call 2: 50.00s. Clearly slow. Price: $3.0 per 1M input and $15.0 per 1M output. Moderately priced for Frontier level, but given this execution stability, not an efficiency case.

Conclusion & Recommendation

Suitable for agentic research and orchestration pipelines where tool selection matters more than perfect result consolidation and where a downstream validator reviews the outputs. Not the first choice for compliance flows, precise MCP automation, or customer-facing direct responses without a control layer. For cloud deployments under Chinese jurisdiction, a clear governance risk applies in addition.

This analysis was generated automatically based on benchmark data. Model used: GPT 4.5 by OpenAI. The raw data and full methodology are documented in the GitHub project.