Alibaba has open-sourced Qwen-Image-2.1, integrating text-to-image and image editing into a single model while balancing quality, efficiency, and cost.
The generation part has 7B parameters, so the parameter count is not large. With everything balanced, it definitely shouldn't be compared with top-tier image models; being usable is enough.
The model's "biggest highlight" is that it supports generating and editing images with an "alpha channel" (whether they are images or text), in simple terms, PNG images finally support transparent backgrounds.
Other features also include:
- Because it supports alpha channels, the model can directly perform cutouts
- Supports inputting 10 reference images at once to synthesize one image, such as combining solo photos of 10 people into a family portrait
- Image editing supports local lasso selection for modifications, and also supports independent masks as input. This enables more precise generation needs; for example, if I want a person to wear a hat of a specified shape, I can input a mask image, so the resulting hat will match the mask
- Use ordinary photos to generate panoramas, with scenes outside the photo extended and generated by the model
- Improved text layout and lighting and shadow detail for people, and strengthened consistency of people and products during editing; this direction is clear and aimed at e-commerce
Overall, in terms of open-source models, Qwen-Image-2.1's capabilities are quite comprehensive. Just writing this makes me want to deploy one for daily use.
