qwen3.7-plus

qwen3.7-plus

Tongyi Qianwen 3.7: The Most Cost-Effective Multimodal Intelligent Agent Foundation in the Family
2026-06-10
LLM
Model capability: imageModel capability: videoModel capability: thinkingModel capability: function_call
Cache Creation:
$0.286/1M tokensstarting from
Cache Read:
$0.023/1M tokensstarting from
Input:
$0.228/1M tokensstarting from
Output:
$0.92/1M tokensstarting from
Bulk order? Contact your manager for exclusive deals

API Overview

qwen3.7-plus is a next-generation multimodal interactive hybrid agent model. Built on the powerful 3.7 text foundation, this model delivers a comprehensive, breakthrough-level upgrade in visual-language capabilities. As a cost-effective solution tailored specifically for complex, high-frequency business scenarios, qwen3.7-plus maintains full-stack agent behavior, tool invocation, and sophisticated planning capabilities that rival those of flagship models while significantly reducing costs. The model seamlessly integrates visual perception with end-to-end system execution, making it the absolute top choice for deploying multimodal automated workflows in highly concurrent production environments.

───────────────────────────────────────────────────────────────────

Core Capabilities

Multimodal Interactive Hybrid Agent: Features groundbreaking cross-modal unified control capabilities. The model can directly perceive real-world scenes, perform high-precision screen reading, and engage in deep, interactive manipulation of complex GUIs. It supports fully autonomous, end-to-end cross-interface navigation and click operations within mobile apps, perfectly integrating visual perception with the intricate interaction loop between underlying CLI and visual elements.

Full-Modal Input and Frontend Prototyping: Supports mixed multimodal inputs including text, images, and videos. It excels in development and productivity scenarios, enabling it to reverse-engineer and faithfully reproduce multilingual underlying code directly from a single design draft or frontend webpage screenshot. It demonstrates high robustness in challenging visual tasks such as automated form parsing, multimodal complex OCR, and fine-grained object localization.

Deep Long-Term Planning and Continuous Thinking: Natively supports the preserve_thinking (continuous thinking preservation) architecture, ensuring that internal thought chains remain uninterrupted and avoid repeated computations during long-term, multi-round interactions in complex engineering tasks. It maintains exceptional trajectory stability in long-term workflows such as multi-step terminal automation and complex hardware optimization.

Ultimate Cost-Effectiveness and High-Concurrency Adaptability: By default, it supports an ultra-long context window of 1 million tokens. Thanks to its exceptionally lightweight architecture and operator acceleration, its first-token latency and throughput both rank among the industry’s best. Without compromising on complex reasoning or strict instruction adherence, it helps enterprises reduce the deployment costs of multimodal agents to around 15% of those for flagship models.


Playground

Log in to explore more features! Click to Log In

API Analytics

API Reference (1)

API DescriptionAPI EndpointRequest MethodStabilityParameter Description
qwen3.7-plus
POST
Stable
View Details

API Pricing

$
ModelDescriptionContextOfficial Price302.AI PriceCache Price

qwen3.7-plus

≤256K tokens
256000

Input$0.285 / 1M tokens
Output$1.15 / 1M tokens

Input$0.228/ 1M tokens
Output$0.92/ 1M tokens
20%

Cache Creation$0.286/ 1M tokens
Cache Read$0.023/ 1M tokens

qwen3.7-plus

>256K tokens
1000000

Input$0.85 / 1M tokens
Output$3.42 / 1M tokens

Input$0.68/ 1M tokens
Output$2.74/ 1M tokens
20%

Cache Creation$0.86/ 1M tokens
Cache Read$0.068/ 1M tokens