Kimi K2.6

Kimi K2.6 is Moonshot AI’s multimodal model for agentic tasks, coding, and tool-assisted workflows, with native input support for text, image, and video. The MoE architecture activates only 32 billion of the total one trillion parameters per token; the context window spans 256,000 tokens. Available as an Open Weights model locally or via cloud API, with Chinese jurisdiction as a material cloud risk factor.

Moonshot AI Version k2.6 Commercial use permitted MoE 1000 B (32 B active) 256 K Context 12/2025 $0.74 / $3.49 per 1M

  • Open Weights
  • Frontier
  • OR
  • Text
  • Vision
  • Video
  • Agentic Orchestrator
  • Long Context
  • Batch

Tool-use profile: 6 assets in detail

Comparison of asset performance (P1/P2/Combined) against the fleet average

Asset performance (radar)

Score Breakdown vs. Fleet Average


Tool-use details

Asset performance, reliability, and runtime profile

CrucibleMark evaluates tool use across 6 independent tests. Click a test name for details.

Reliability

  • Tool Call Valid: No
  • Retry: Not required
  • Hallucination: Not detected

Reliability measures how consistently a model actually executes tool calls: Tool Call Valid schema and format accepted, Retry Required successful only after retry, Hallucination Flag fabricated tools or parameters detected. All three green means production-ready.

Runtime profile

Call 1
9.77
First Request
MCP
1.54
Protocol Latency
Synthesis
26.39
Response Generation
Total
226.23
Sum of All Phases
Token
5386
Input + Output
Cost
$0.0089
Cost per Run

The runtime profile shows the latency and cost metrics for the model run: Call 1 First Request, MCP Protocol Latency, Synthesis Response Generation, Total sum of all phases, supplemented by Token input and output and Cost cost per run.

Tool-use review

· Agentic Orchestrator · Long Context

Deployment Verdict

Created on: 14.06.2026, 16:13:36

Conditional deploy, because Kimi K2.6 executes tool calls validly and without hallucination, but synthesis quality at Combined 74.50 is not stable enough for high-stakes output pipelines.

Tool Execution Profile

The model is reliable on the execution side. Tool call valid, no retry required, no protocol anomalies. This speaks to clean MCP integration in production workflows.

For Web Search & Tool Selection — the test of whether the appropriate research tool is chosen without a hint — it scores P1 80. For the URL Construction test, which checks autonomous derivation of a target URL and the subsequent fetch, it also lands at P1 80. This shows no deep tool intelligence, but no rigid failure pattern either. Kimi K2.6 recognizes the fundamental difference between a search step and a direct retrieval and executes both paths serviceably. For deterministic pipelines, however, some degree of oversight remains necessary, as the selection is correct enough but not precise enough for blind pass-through.

Synthesis Fidelity

How well does it condense tool results? Only adequately. P2 63.33 is the clear bottleneck of this model. The pattern holds consistently across assets: retrieval succeeds, but condensation loses precision, nuance, or prioritization. For simple summaries, this is sufficient. For compliance, policy excerpts, or decision-relevant extraction, post-review is required.

Does it stay within the tool result or fall back on training? The signal here is good. In the EU License Research test — a honeypot test for current license restrictions from web sources — the model stayed within the retrieved material. Content Verification State A and no detected hallucination are the actual trust anchor of this run. It does not answer from implicit prior knowledge when fresh sources are required.

Error Resilience

For Tool Failure Handling with 404 — the test for transparent handling of failed retrievals — Kimi K2.6 responds in a production-appropriate manner. P2 80 combined with no hallucination despite a 404 is a good signal. The model communicates the error state rather than fabricating missing page content. For real-world tool pipelines, this is acceptable.

Operational Profile

Call 1: 9.77s. Call 2: 26.39s. MCP latency: 1.54s. Total: 226.23s. Slow overall. Cost per run: 0.008944 USD. Inexpensive to moderate relative to the performance demonstrated.

Conclusion & Recommendation

Suitable for agentic research and orchestration pipelines in which the model reliably invokes tools, reports errors transparently, and a downstream validator reviews the condensation. Not the right choice for pipelines where the initial textual synthesis must already be decision-ready. If you deploy Kimi K2.6, do so as a tool operator with a controlled output stage — not as an unsupervised final author.

This analysis was generated automatically based on benchmark data. Model used: GPT 4.5 by OpenAI. The raw data and full methodology are documented in the GitHub project.