Alibaba has launched its latest foundational model for image generation, Qwen-Image-3.0. This model supports exceptionally long inputs of up to 4.5k tokens, enabling the generation of complex, multi-element visuals in a single pass. These can include knowledge diagrams featuring formulas, symbols, geometric shapes, and logical deduction steps, as well as intricate UI interfaces. It natively supports rendering in 12 languages and over 20 fonts, significantly reducing the production cost for commercially viable materials like multilingual product posters and film/comic storyboards.
Qwen-Image-3.0 represents the third generation in the Qwen series of image generation and editing models. It can effortlessly generate content in areas where predecessor models excelled, such as live-streaming room layouts. The model also produces highly detailed and lifelike images in categories like portrait photography, math test papers, and ancient stone rubbings.
Key Model Enhancements
The new model places a strong emphasis on enhancing its capacity to handle complex information, achieving a 4.5-fold increase in text input length compared to its predecessor. This advancement is particularly beneficial for scenarios requiring extensive prompts, such as creating storyboards for short videos, complex infographics, or product specification pages. Users no longer need to compress their requirements into a single sentence. Instead, they can provide a comprehensive description akin to a design brief, detailing the image structure, text content, visual style, and layout specifics. This allows the model to generate the desired visual content more accurately, reducing the need for repeated generation attempts and manual adjustments.
Advanced Capabilities in Generation
Leveraging Qwen-Image-3.0's refined control over semantic juxtaposition and spatial relationships, the model can generate a nine-panel knowledge diagram covering nine different domains in one go, with each panel displaying clear text and accurate imagery.
Furthermore, Qwen-Image-3.0 excels at understanding complex instructions and the logical nesting relationships within a scene. It can generate a composite visual containing multiple types of content—such as a webpage, a software interface, a chat window, and a poster—all within a single image based on one instruction. For example, when prompted to "generate a pour-over coffee poster published in a chat interface, which is inside the Qwen App, which is itself within a VSCode programming interface," the model accurately interprets the multi-layered UI nesting. It sequentially generates the VSCode interface, the Qwen App interface, the social media chat window, and finally the coffee poster. It maintains the structural and stylistic accuracy of each layer, with details as fine as the time displayed on a phone screenshot and tiny text as small as 10 pixels remaining crisp and legible.