AI's ability to write code often allows developers to "take their hands off the keyboard," but for designers, AI-generated images are still several steps away from being deliverable.
In the image design workflow, elements such as人物 need to be cut out, text on assets needs to be modified, and the placement of various products requires occasional adjustments—however, the text on the packaging and the shape of the bottle must remain unchanged. If requirements continue to evolve, localized edits must still be made to ensure the entire image remains cohesive.
These details determine whether AI-generated images can enter the true creative workflow. For designers, e-commerce operators, and content creators, the current bottleneck in image models lies in whether each modification instruction can be accurately executed.
This Sunday, Alibaba Qwen Officially open-sourcing the image generation model Qwen-Image-2.1, which addresses most productivity challenges with a compact model size: it integrates text-to-image and image editing, natively supports transparent image generation, accepts up to 10 reference images, and further enhances local editing, portrait fidelity, and product realism.

It became Qwen The Qwen-Image-2.1 is currently the most balanced and cost-effective open-source image generation model. Public evaluation results show that Qwen-Image-2.1 ranks first among open-source models, even surpassing Nano Banana 2.0. Notably, the visual generation component of this model has only 7B parameters.
On platforms like X, users who received early access have shared their generated results, receiving widespread praise. There are renderings containing large amounts of text:

Generate images with the texture of real photographs:

There have also been attempts with various filters and lighting styles:

Qwen-Image-2.1 is widely regarded as impressive in instruction following, draw success rates, and real-world performance. Leveraging the capabilities of this new model, we can now create seamless workflows for asset generation, composition, and editing, using AI to address many complex image editing challenges.
Capability and Efficiency
According to the information provided, the visual generation component of Qwen-Image-2.1 employs a 20-layer Single-Stream DiT with 7 billion parameters and natively supports 2K resolution. The model achieves performance beyond its size on evaluation benchmarks: Qwen-Image-2.1 scores a total of 60.28, compared to 59.82 for Nano Banana 2.0 and 59.65 for GPT Image 1.5. It is now competitive with several proprietary image models.

Qwen-Image-Bench evaluation chart.
The visual generation component has only 7B parameters, making its capabilities even more impressive: it integrates image generation, transparent asset creation, and multi-image editing within a smaller generation module (unified generation and editing pipeline), offering developers a lighter-weight starting point for deploying and customizing image tools.
While delivering strong performance, Qwen-Image-2.1 also optimizes the model's inference process by reducing unnecessary repeated computations.
During image editing, the input reference image and editing instructions remain unchanged across each generation step. Therefore, the new model employs a hybrid-granularity attention architecture: the text components (system prefix, editing instructions) follow traditional token-level causal masking, performing strict sequential computation word by word to ensure coherent semantic understanding, while the image generation component uses chunk-level masking and leverages the KV Cache mechanism to cache these static contexts after the initial computation, allowing them to be reused in subsequent generation steps.

This means AI will now process reference materials first and reuse existing results when gradually generating target images. This optimization is especially effective with multiple image inputs, helping to improve inference efficiency and reduce VRAM usage.
Since Qwen-Image-2.1 now supports up to 10 reference images, this optimization has become crucial. The richer the input, the more important it is to effectively manage reference information and control redundant computations. The exact amount of time and VRAM saved will depend on factors such as hardware, precision, resolution, and the number of inputs.
Directly generate usable transparent assets
The capability of the new Qwen-Image version that best aligns with design requirements is transparent image generation.
Regular image assets typically come with complete backgrounds. To use the main subject in posters, product detail pages, or other images, you often need to first cut out the subject and refine its contours, gaps, and edges. Transparent images include additional alpha channel information, allowing the background to show through around or within parts of the subject, making it easier to layer and combine later.
Qwen-Image-2.1 natively supports transparent image generation and can automatically determine whether to generate a standard image or a transparent image based on the prompt. Users can request a transparent background directly in their description of the asset. The current examples include holiday illustrations, character assets, decorative elements, and complex compositions made up of multiple objects—all common asset requirements in design workflows. We also tried:

This is great news for creators—we can now access ready-to-use assets, eliminating several steps from our workflow.
Of course, these transparent assets can be further edited after generation. For example, the same illustration character can have their expression adjusted, and the text in the image can be modified. The editable elements can include visual details of the character or the text within the asset.
This closely matches the actual modification requirements during design: the event name has changed, but the artwork must be retained; the expression should be more relaxed, while the character must still be recognizable as the same person. After combining generation and editing into a single model, these subsequent requirements now have a unified processing interface.
Of course, Qwen-Image-2.1 supports processing existing photos.


The official documentation provides examples of extracting required content from real photos to obtain RGBA transparent images, where "A" represents the alpha channel, which describes the level of transparency at each point in the image.
Clearly, this capability supports both generating content from scratch and repurposing existing images. Creators can first upload a photo and then request the model to extract the desired subject as素材 for subsequent design.
The Qwen-Image series previously introduced the standalone Qwen-Image-Layered to explore capabilities related to transparent images. With Qwen-Image-2.1, native transparent image generation and editing have been integrated into the general-purpose image model, further enhancing convenience.
Accurate, open-source AI models for multi-image editing are now available.
When an image's requirements come from multiple sources, the editing task becomes even more complex.
For example, if you want to dress the same model with a specified outfit, shoes, bag, and hat, each reference image serves a different purpose, so the model must distinguish between the person and the items, place them in appropriate positions, and preserve their individual features as much as possible.
Qwen-Image-2.1 supports up to 10 reference images. We tried it out, having Ultraman and Dario sit down for a conversation. We used seven images: a photo of the two facing each other (with adjusted expressions), the background room, a single-seater sofa, a coffee cup (filled with coffee), a floor lamp, an electric fireplace, and a low table, ultimately generating a complete rendered image.

You can now quickly visualize how different outfits look on a person, and you can also combine individual photos into group photos.
Technically, they test not only how many images an AI model can process as input, but also whether the reference objects are properly positioned, proportioned, and stylistically consistent in the final output.
The overall image has been composed; typically, further fine-tuning follows.
Qwen-Image-2.1 provides options such as selection, brushing, and independent masking to help users more clearly specify editing areas. For example, you can use multiple colors to mark different regions and request a single edit to simultaneously remove clothing from a person in the photo and change their hair color.

This interaction establishes a correspondence between spatial locations and text requirements. For edits involving multiple areas, users can specify exactly where to delete, where to replace, and what to change it to.
The brushing method is suitable for directly marking the location of new objects, but circling or brushing directly on the original image may obscure part of the画面. To address this, the model also supports specifying the editing area using an additional mask image, allowing the full original image information to be fed into the model. This ability to preserve original details is especially critical for designs containing people and products.

Text accuracy, visual quality
Even if the hairstyle or scene changes, the portrait should retain its original identifying features; when a product is placed in a different environment, its packaging text, texture, and shape should remain as consistent as possible. Otherwise, even if the image looks appealing, it may not be suitable for its intended display purpose. Qwen-Image-2.1 highlights its fidelity in editing both portraits and products, and the results appear quite promising.
We tried it, for example, by having it process a screenshot from this page of the primary school textbook "Information Technology, Grade 3, Semester 1": DeepSeek The welcome message: "I am DeepSeek Hello, how can I help you? Update the Deep Thinking (R1) button to the latest version and activate it:

Beyond generation and editing, Qwen-Image-2.1 has further improved its performance on text and human figures. Whether in text-dense image pages, infographics, or illustrations for academic papers, the accuracy of content and the aesthetic quality of layout have been significantly enhanced.

In the portrait section, Qwen-Image-2.1 has further enhanced light and shadow effects and detail rendering. AI-generated images now exhibit more harmonious representation of hair strands, fabric textures, facial details, and ambient lighting, with significantly improved visual texture.
Additionally, it offers enhanced capabilities for generating panoramic images, infographics, three-view diagrams, and storyboards, expanding the applications of static image generation and editing, and providing more possibilities for character presentation, visual planning, and similar tasks. We tried using community-created Qwen anime-style characters to generate a series of storyboard illustrations:

Together, these capabilities enable Qwen-Image-2.1 to cover a more comprehensive image processing workflow—allowing us to accomplish extensive tasks using a single AI model: first organizing the composition with multiple reference images, then adjusting specific regions while preserving key features of people and products. How much this open-source AI model, with only 7B parameters in its visual generation component, can reduce revision cycles will be an important question to evaluate in subsequent real-world testing.
How large is the application space?
The release of Qwen-Image-2.1 shows that image models are increasingly covering more detailed stages of the creative process, and this trend is accelerating.
Whether it’s the output of transparent assets, multi-image reference, localized editing, or improved fidelity, these advancements increase the likelihood of AI models being integrated into real-world workflows. For creators, this means more asset creation and adjustment tasks can be handled more simply through open-source image generation models; for developers, the open-sourcing of these new models provides a foundational basis for building design tools, applications, and customized workflows around these capabilities.
We can boldly predict that, with the open-sourcing of image models like Qwen-Image-2.1, the community will further integrate generation and editing capabilities into more content creation scenarios, enabling more people to access creative tools tailored to their needs. When these AI capabilities become routinely available features in everyday software, the barrier between idea and visual creation may once again be significantly lowered.
The open weights for Qwen-Image-2.1 have been released on Hugging Face and ModelScope, and the technical report has been uploaded to the GitHub repository.
Blog: https://qwen.ai/blog?id=qwen-image-2.1
GitHub: https://github.com/QwenLM/Qwen-Image-2.1
Model Scope: https://www.modelscope.cn/models/Qwen/Qwen-Image-2.1
Hugging Face: https://huggingface.co/Qwen/Qwen-Image-2.1
This article is from the WeChat public account "Machine Heart" (ID: almosthuman2014), authored by Zeanan.
