deepseek-v4.1-flash

DeepSeek-V4.1-Flash is the lightweight flagship of DeepSeek's new architecture family, packing 552B total MoE parameters to deliver flagship-surpassing intelligence across key benchmarks, including out-performing DeepSeek-V4-Pro. It adopts a Causal Encoder-Decoder asymmetric design with 8B active parameters for input and 16B for output, and offers native multimodal visual understanding. KV Cache usage is cut to one-quarter of the previous generation's HBM and one-eighth of its SSD storage, drama

LiveDeepSeek3 protocols1M contextStream cancellation unsupported
Context
1M
Input / 1M
$0.30
Output / 1M
$1.20
Modalities
Protocol
3
  • chat.completions
  • messages
  • responses

Status & performance

Last 7 days
Loading model performance

Capabilities

TextVisionLong contextCache

Pricing

Time-based pricing

Weekly schedule · Asia/Shanghai

When matched, the complete price for that range replaces the default; missing items are not filled from it.

00:0006:0012:0018:0024:00

Default pricing

Applies when no time range matches on Every day
Input tokens
$0.3/1M
Output tokens
$1.20/1M
Cache read
$0.03/1M
Cached input
$0.03/1M

Prices in USD per 1M tokens unless noted otherwise. Batch calls and cache hits receive additional discounts; live rates apply in the Workspace.