kimi-k3

kimi-k3

The new-generation 2.8 trillion-parameter multimodal large model of Moonshot
2026-07-17
LLM
Model capability: imageModel capability: videoModel capability: thinkingModel capability: function_call
Cache Read:
$0.3/1M tokens
Input:
$3/1M tokens
Output:
$15/1M tokens
Bulk order? Contact your manager for exclusive deals

API Overview

Kimi K3 is a massive 2.8-trillion-parameter multimodal reasoning large model—part of a new-generation series designed specifically for agent programming and fully autonomous knowledge workflows, tailored for the far side of the moon. As a game-changer in the global open-source landscape, K3 leverages an innovative Kimi incremental attention mechanism and a hybrid sparse expert architecture, achieving a 2.5-fold increase in overall scaling efficiency compared to its predecessor, K2. Its performance in software engineering delivery, multimodal visual reasoning (such as frontend/CAD code generation), and large-scale swarm agent collaboration has propelled it firmly into the global top three among large models, making it the undisputed king for enterprise-level complex scenarios.

───────────────────────────────────────────────────────────────────

Core Capabilities

2.8 Trillion Parameters: At the Pinnacle of Global Open Source: K3 sets a new benchmark for scale and intelligence among global open-source large models, becoming the first ultra-large-scale reasoning foundation to surpass the 2-trillion-parameter threshold. Thanks to its Stable LatentMoE architecture, which activates 16 experts out of 896, the model not only demonstrates cutting-edge “high-IQ” logical reasoning on a global scale but also maintains exceptional stability in inference throughput.

Fully Enabled Adaptive Deep Thinking Mode: By default, K3 natively enables deep reasoning chains (CoT) without requiring any manual intervention, allowing it to reflect upon and correct complex boundary conditions automatically. It supports fine-grained configuration of reasoning effort (Reasoning Effort; max level by default). When dealing with high-complexity financial audits, highly rigorous legal document flows, and intricate mathematical logic, K3 exhibits an exceptionally high “first-attempt pass rate.”

Long-Range Coding and Visual Reasoning Closed Loop: K3 has been rigorously optimized for complex automated software engineering tasks. With minimal human intervention, K3 can run autonomously and continuously, comprehending tens of thousands of lines in large private codebases and using terminal tools to close the loop and resolve real-world project vulnerabilities. Additionally, it supports native visual alignment, enabling fully automatic optimization of 3D game development and frontend layout through screen captures and runtime feedback.

Efficient Collaboration Among Swarm Agents: K3 is perfectly adapted to next-generation swarm agent clusters and goal-oriented parallel workflow patterns. Within a context window of up to one million tokens, K3 can serve as the supreme commander, coordinating and orchestrating massive numbers of agents executing tasks in parallel. Its needle-in-a-haystack recall accuracy covers the entire span flawlessly, with no information fragmentation or instruction deviation whatsoever.

Playground

Log in to explore more features! Click to Log In

API Analytics

API Reference (1)

API DescriptionAPI EndpointRequest MethodStabilityParameter Description
Chat(Moonshot kimi AI-Vision)
POST
Stable
View Details

API Pricing

$
ModelDescriptionContextOfficial Price302.AI PriceOfficial Price Gap

kimi-k3

-
1000000

Input$3 / 1M tokens
Output$15 / 1M tokens

Cache Read$0.3 / 1M tokens
Input$3 / 1M tokens
Output$15 / 1M tokens

Original Price