Kimi K3

Kimi K3 has been Moonshot AI’s Frontier flagship since July 2026, with 2.8 trillion total parameters making it the largest Open Weights model in the world to date. The Stable-LatentMoE architecture activates only 16 billion parameters per token from 896 experts; the context window spans 1,000,000 tokens for long-horizon agentic workflows. Kimi Delta Attention and Attention Residuals are designed to significantly improve scaling efficiency over the predecessor model.

Moonshot AI Version k3 Commercial use permitted MoE 2800 B (16 B active) 1000 K Context 02/2026 $3 / $15 per 1M

  • Open Weights
  • Frontier
  • OpenRouter
  • Text
  • Vision
  • Video
  • Long Context
  • Agentic Orchestrator
  • Batch

Sovereign Risk: HIGH Moonshot AI is a Chinese company subject to China’s National Security Law (NSL), which may enable state access to data via services hosted in China. In February 2025, Germany’s BSI explicitly warned against the use of Chinese AI cloud services; this risk assessment applies here conservatively as well. Kimi K3 has been announced as a 2.8T Open Weights model whose weights are scheduled for release on July 27, 2026; the operational risk arises primarily from cloud usage under Chinese jurisdiction, not from the mere existence of open weights.[web:605][web:610]

Key metrics

Score · Latency · Cost · Quality

Total Score Silver
78.83
Routine
47.71
Reasoning
31.12

Rank #6

LLM Judge Avg
3.95
100 Coverage
Avg Task Duration
99.3
Batch
Token Rate
48.28
Output Rate
P95 Latency
310.8
Top 5 %
Total Tokens
219900
Output Volume
Cost per 1K
$0.015
USD / 1K Requests
Benchmark Cost
$3.3
Total · 219900 tok

Benchmark modules

10 modules · weighted · vs. model median & top performer

Kimi K3 Best model Ø All models
Code Quality 78.64
CLI Benchmark 90.67
Logical Reasoning 75.27
UX Writing 82.31
Documentation 78.33
Content Transform. 75.87
Cultural Intelligence 74.64
Synthesis Quality 73.33
Tool Execution 88.33
ToolUse Score 80.83
Benchmark Cost $3.3

Token efficiency & latency

Consumption per module vs. model median

Token consumption per module

Performance profile