Kimi K2.7 Code

Kimi K2.7 Code is the coding-specialized agentic model from Moonshot AI within the K2 family, optimized for long-horizon software engineering workflows. The MoE architecture activates 32 billion out of a total of one trillion parameters per token, with a context window of 256,000 tokens. Multimodal input for text, image, and video, always-on thinking mode, and tool use support. Available as an Open Weights model under a Modified MIT license.

Moonshot AI Version 2.7-code Commercial use permitted MoE 1000 B (32 B active) 256 K Context 10/2025 $0.67 / $3.4 per 1M

  • Open Weights
  • Frontier
  • OpenRouter
  • Text
  • Vision
  • Video
  • Instruction-Tuned
  • Agentic Orchestrator
  • Batch

Sovereign Risk: HIGH Moonshot AI is a Chinese company and subject to China’s National Security Law (NSL), which may enable state access to data. In February 2025, the BSI explicitly warned against the use of Chinese AI cloud services; this risk assessment conservatively applies here as well.

Tool-use profile: 6 assets in detail

Comparison of asset performance (P1/P2/Combined) against the fleet average

Asset performance (radar)

Score Breakdown vs. Fleet Average


Tool-use details

Asset performance, reliability, and runtime profile

CrucibleMark evaluates tool use across 6 independent tests. Click a test name for details.

Reliability

  • Tool Call Valid: No
  • Retry: Not required
  • Hallucination: Detected

Reliability measures how consistently a model actually executes tool calls: Tool Call Valid schema and format accepted, Retry Required successful only after retry, Hallucination Flag fabricated tools or parameters detected. All three green means production-ready.

Runtime profile

Call 1
4.03
First Request
MCP
1.13
Protocol Latency
Synthesis
12.72
Response Generation
Total
107.33
Sum of All Phases
Token
11908
Input + Output
Cost
$0
Cost per Run

The runtime profile shows the latency and cost metrics for the model run: Call 1 First Request, MCP Protocol Latency, Synthesis Response Generation, Total sum of all phases, supplemented by Token input and output and Cost cost per run.

Tool-use review

· Instruction-Tuned · Agentic Orchestrator

Deployment Verdict

Conditional deploy: The model makes tool decisions correctly in many cases, but is not reliable enough for production MCP pipelines without guardrails — hallucination was detected, tool calls were not consistently valid, and overall yield remains only moderate.

Tool Execution Profile

Kimi K2.7 Code demonstrates genuine tool intelligence rather than pure schema-following. On the Web Search & Tool Selection test — which checks the choice between search and direct fetch without an explicit hint — it identifies the need for web_search with high confidence. This speaks to usable orchestration in open research paths. On the URL Construction test, which requires deriving the target URL from internal knowledge and then fetching it correctly, it is less precise. P1 80 is workable, but not strong enough for deterministic pipelines. More critically, Tool-Call valid overall reads false. This means the model is not MCP-safe in the strict sense. It understands the workflow but does not consistently produce protocol-clean execution. The absence of any retry needed points less toward a pure formatting issue and more toward inconsistent first-attempt execution.

Synthesis Fidelity

How well does it consolidate tool results? Only to a limited degree. P2 70 looks acceptable at first glance, but the spread across assets is too wide for production use. HTTP Fetch & Extract drops off noticeably in consolidation, and Multilingual Search & Synthesis as well as EU License Research land effectively at zero for result consolidation. The model can use tools, but loses precision when translating results into reliable final statements.

Does it stay within tool output or fall back on training data? The honeypot signal is negative, even without a formal hallucination in any individual case. On the EU License Research test — which checks whether current license restrictions are answered from web sources rather than parametric prior knowledge — it delivers P2 0. That is a trust problem. In addition, hallucination has been detected globally. In a tool pipeline, this is not merely a quality deficiency but a security risk: the model can output invented or insufficiently substantiated claims framed as tool results.

Error Resilience

On the 404 test, which measures transparent behavior when a tool call fails, the model responds acceptably. It communicates the failure in an essentially open manner and does not fabricate page content. P2 80 and no hallucination despite a 404 are a positive signal for production. At the infrastructure level, it does not reflexively fall back on substitute facts.

Operational Profile

Total 107.33s per run. Slow. Individual calls 4.03s and 12.72s, MCP latency 1.13s. Cost/run local. Model price level: $0.67 per 1M input, $3.4 per 1M output. For the reliability demonstrated, the operational profile is more of a burden than an efficiency.

Conclusion & Recommendation

Suitable for internal engineering assistance with downstream validation, particularly where tool selection matters more than precise final consolidation. Not suitable for compliance, research, or approval pipelines in which tool results must be merged without distortion. If you deploy it, do so only with strict tool-call validation, mandatory source attribution per statement, and a second instance for result verification before output.

This analysis was generated automatically based on benchmark data. Model used: GPT-5.4 by OpenAI. The raw data and full methodology are documented in the GitHub project.