qwen3.8-max

qwen3.8-max

Alibaba’s brand-new 2.4 trillion-parameter flagship large model
2026-08-04
LLM
Model capability: imageModel capability: videoModel capability: thinkingModel capability: function_call
Cache Creation:
$2.15/1M tokens
Cache Read:
$0.14/1M tokens
Input:
$1.8/1M tokens
Output:
$5.3/1M tokens
Bulk order? Contact your manager for exclusive deals

API Overview

Qwen3.8-Max is Alibaba’s next-generation flagship large model, boasting 2.4 trillion parameters and specially designed for long-term autonomous agents, complex multimodal reasoning, and system-level control. As the pinnacle of the Qwen3.8 family, this model deeply integrates a unified vision-language-code reasoning architecture at its core, breaking away from the traditional AI limitation of responding only to single prompts. The model exhibits exceptional self-planning and self-reflection capabilities, enabling it to autonomously carry out large-scale software engineering projects, automated replication of academic papers, and iterative complex chip designs over several consecutive days. It has firmly topped the Agentic Computer Use benchmark.

───────────────────────────────────────────────────────────────────

Core Capabilities

2.4 Trillion Parameter Long-Term Agent Dominance: The model achieves a breakthrough in industrial-grade long-task execution. It can completely autonomously decompose complex automated engineering projects spanning several days. When encountering errors along the way or facing unexpected UI obstacles, it demonstrates remarkable resilience through “self-reflection and trajectory replanning,” completely eliminating deadlocks and infinite loops.

Superior Computer Use Visual Control: The model has been deeply optimized for GUI interactions with operating systems and multimodal screen perception. Like a human expert, it can precisely identify high-resolution desktops, professional software UI elements, and intricate CAD/chip design blueprints, automatically executing highly challenging cross-software chain operations.

Hundred-Word-Level Thought Chains and 131K Ultra-Long Outputs: The model natively supports deep-thinking modes and offers multi-level adjustment options for thinking intensity—low, medium, high, and max. Not only does the model natively support ultra-large context inputs of up to 1 million tokens, but its single-generation output limit has also been expanded to an astonishing 131,072 tokens, allowing it to fully generate massive codebases or lengthy, in-depth research reports in one go.

End-to-End Code Replication and Self-Developed Research: Specifically designed for intensive scientific research and cutting-edge engineering applications. Without any human intervention, the model can autonomously write tens of thousands of lines of engineering code and complete closed-loop verification solely based on an academic paper PDF or design documentation, demonstrating comprehensive delivery capabilities comparable to those of top-tier senior architects.

Playground

Log in to explore more features! Click to Log In

API Analytics

API Reference (1)

API DescriptionAPI EndpointRequest MethodStabilityParameter Description
qwen3.8-max
POST
Stable
View Details

API Pricing

$
ModelDescriptionContextOfficial Price302.AI PriceOfficial Price Gap

qwen3.8-max

-
1000000

Input$2 / 1M tokens
Output$6 / 1M tokens

Cache Creation$2.15 / 1M tokens
Cache Read$0.14 / 1M tokens
Input$1.8 / 1M tokens
Output$5.3 / 1M tokens

Original Price