qwen3.8-flash

Qwen3.8-Flash is the latest multimodal model from the Qwen family, combining powerful reasoning and generation with remarkable speed. It natively supports a million-token context window, allowing it to process lengthy documents, entire codebases, and complex conversations in a single pass. It shines in coding assistance, agentic workflows, and visual understanding — whether it's fixing code autonomously, operating desktop applications, or analyzing charts and long videos. Fully compatible with b

LiveAlibaba3 protocols1M contextStream cancellation unsupported
Context
1M
Input / 1M
$0.15
Output / 1M
$0.47
Modalities
Protocol
3
  • chat.completions
  • messages
  • responses

Status & performance

Last 7 days
Loading model performance

Capabilities

TextVisionLong contextCache

Pricing

Input tokens
$0.15/1M
Output tokens
$0.47/1M
Cache read
$0.016/1M
Cached input
$0.016/1M
Cache write (5m)
$0.2/1M
Cache write (1h)
$0.2/1M
Cache write
$0.2/1M

Prices in USD per 1M tokens unless noted otherwise. Batch calls and cache hits receive additional discounts; live rates apply in the Workspace.