deepseek-v4.1-flash
DeepSeek-V4.1-Flash is the lightweight flagship of DeepSeek's new architecture family, packing 552B total MoE parameters to deliver flagship-surpassing intelligence across key benchmarks, including out-performing DeepSeek-V4-Pro. It adopts a Causal Encoder-Decoder asymmetric design with 8B active parameters for input and 16B for output, and offers native multimodal visual understanding. KV Cache usage is cut to one-quarter of the previous generation's HBM and one-eighth of its SSD storage, drama
- chat.completions
- messages
- responses
Status & performance
Capabilities
Pricing
Time-based pricing
Weekly schedule · Asia/ShanghaiWhen matched, the complete price for that range replaces the default; missing items are not filled from it.
Default pricing
Applies when no time range matches on Every dayPrices in USD per 1M tokens unless noted otherwise. Batch calls and cache hits receive additional discounts; live rates apply in the Workspace.
