qwen3.8-flash

qwen3.8-flash

Alibaba’s brand-new, cost-effective multimodal large model
2026-09-01
LLM
Model capability: imageModel capability: videoModel capability: thinkingModel capability: function_call
Cache Creation:
$0.2/1M tokens
Cache Read:
$0.016/1M tokens
Input:
$0.15/1M tokens
Output:
$0.47/1M tokens
Bulk order? Contact your manager for exclusive deals

API Overview

Qwen3.8-Flash is Alibaba’s next-generation, ultra-cost-effective multimodal large model designed for high concurrency, complex long-range agents, and engineering-grade code delivery. At the architectural level, the model achieves a groundbreaking breakthrough by integrating a hybrid mechanism that combines Gated DeltaNet with Qwen Sparse Attention (QSA). It also pioneers an innovative approach of asynchronously offloading a 51B-parameter N-gram embedding table to memory offline. As the latest benchmark in the Flash series, this model not only accurately handles ultra-long contexts up to 1M tokens and high-density visual/UI scenes but also significantly boosts long-sequence generation speed with remarkably low computational overhead—making it the ideal choice for enterprises scaling AI code assistants, agent clusters, and visual workflows.

───────────────────────────────────────────────────────────────────

Core Capabilities

Qwen4 Innovative Architecture Preview: Adopting an ultra-sparse MoE architecture (total parameters: 125B; only 6B activated per token). By combining QSA micro-block-level sparse attention with a 51B-memory-offloaded N-gram embedding, the model boosts long-context Prefill speed by up to 7.6 times, dramatically reducing computational and inference costs.

Code Engineering and Agent Dominance: Specifically designed for complex software refactoring and automated agent workflow optimization. The model demonstrates outstanding performance in real-world repository-level repair evaluations such as SWE-bench Pro, exhibiting exceptional resilience in its “planning-execution-tool invocation-self-debugging” closed-loop capability.

Multi-Level Controllable Reasoning and Adaptive CoT: Native support for deep logical reasoning, offering multiple levels of inference intensity—from high to low. Developers can dynamically balance reasoning IQ and first-token response latency according to specific application scenarios.

Native High-Density Visual and UI Interaction: A single API seamlessly supports text, high-resolution charts, and dynamic video analysis. The model delivers extremely high precision in mobile UI automation testing and complex visual mathematical reasoning, enabling it to faithfully convert design drafts and error screenshots into engineering code.


Playground

Log in to explore more features! Click to Log In

API Analytics

API Reference (1)

API DescriptionAPI EndpointRequest MethodStabilityParameter Description
qwen3.8-flash
POST
Stable
View Details

API Pricing

$
ModelDescriptionContextOfficial Price302.AI PriceOfficial Price Gap

qwen3.8-flash

-
1000000

Input$0.15 / 1M tokens
Output$0.47 / 1M tokens

Cache Creation$0.2 / 1M tokens
Cache Read$0.016 / 1M tokens
Input$0.15 / 1M tokens
Output$0.47 / 1M tokens

Original Price