kimi-k3

kimi-k3 is Kimi's flagship model, built for long-horizon coding and end-to-end knowledge work, with native visual understanding. K3 always runs in thinking mode. You can set the thinking effort via the top-level reasoning_effort parameter (supports low, high, max; defaults to max). It supports a context window of up to 1M tokens.

LiveMoonshot2 protocols1M contextStream cancellation unsupported
Context
1M
Input / 1M
$3.00
Output / 1M
$15.00
Modalities
Protocol
2
  • chat.completions
  • messages

Status & performance

Last 7 days
Loading model performance

Capabilities

TextVisionLong contextCache

Pricing

Input tokens
$3.00/1M
Output tokens
$15.00/1M
Cache read
$0.3/1M
Cached input
$0.3/1M

Prices in USD per 1M tokens unless noted otherwise. Batch calls and cache hits receive additional discounts; live rates apply in the Workspace.