
qwen-image-3.0-pro
API Overview
Qwen-Image-3.0-Pro is the flagship foundation model of the Qwen model family, representing a new generation of visual image generation and retouching/editing capabilities. As the Pro flagship version of the Qwen-Image 3.0 series, this model moves away from the traditional approach of simply enhancing aesthetic appeal and boosting image quality in text-to-image models, instead focusing on high-density information presentation and practical productivity applications. The model not only boasts exceptional world knowledge and native text rendering capabilities in 12 languages, but also natively handles multi-layered nested UI prototypes, as well as ultra-high-definition visual creations at 2K/4K resolution—including complex mathematical formulas and newspaper layouts. It serves as an industrial-grade image-generation solution for e-commerce operations, design layout, instructional illustrations, and advertising marketing.
───────────────────────────────────────────────────────────────────
Core Capabilities
Ultra-long Prompt Adherence: Traditional image-generation models tend to lose track of prompts exceeding dozens of words, whereas Qwen-Image-3.0-Pro natively supports complex long-form texts and lengthy briefs of up to 4,500 tokens. With just one long prompt, you can generate—on a single canvas—a 3×3 grid of long comics, infographics, instructional flowcharts, or even full-page newspapers all at once, eliminating the need to generate pieces individually and then stitch them together later.
Microscopic Text and Precise Arithmetic Rendering: This model completely addresses the common pain points of traditional AI image generation, such as garbled text and distorted symbols. It supports native typesetting in 12 languages including Chinese, English, Japanese, Korean, Spanish, and more. Even when font sizes are as small as 10 pixels (10px), the text remains perfectly legible and highly faithful, capable of rendering academic LaTeX formulas with superscripts, subscripts, and Greek letters in high fidelity.
Multi-layer Nested UI and Realistic Scene Simulation: The model can perfectly simulate complex UI prototypes of mainstream websites, games, and apps within a single image. It even effortlessly handles four-layer visual penetration layouts—such as embedding chat windows inside IDE interfaces, or embedding posters inside chat windows—while ensuring that the styles of each UI layer remain completely independent and do not interfere with one another.
2K Native Generation and Localized Precision Retouching/Editing: The model natively supports high-resolution image outputs up to 2048×2048 pixels, vividly reproducing details like skin pores, hair strands, and fabric textures. At the same time, it allows users to input up to three reference images for style reshaping, localized additions or deletions, restoration of ancient paintings, and face-swapping or scene-changing while preserving facial features.
API Console
Log in to explore more features! Click to Log In