Qwen multimodal models developed by Alibaba Cloud offer advanced capabilities in image and video generation. These models excel at creating high-quality visual content from text descriptions with a strong understanding of both Chinese and English prompts.
Qwen Image 3.0 — Standard & Pro Generation and Editing
- alibaba/qwen-image-3.0/text-to-image: High-quality text-to-image generation model with strong instruction understanding, detailed rendering, coherent compositions, and flexible creative control.
- alibaba/qwen-image-3.0-pro/text-to-image: Professional-grade text-to-image model with enhanced detail, advanced prompt understanding, superior visual quality, and up to 2K output.
- alibaba/qwen-image-3.0/edit: High-quality image editing model for transforming existing images with natural-language instructions while preserving visual consistency and subject identity.
- alibaba/qwen-image-3.0-pro/edit: Advanced image editing model with stronger instruction understanding, superior visual quality, precise control, and up to 2K output for professional workflows.
Qwen Image 2.0 — Standard & Pro Generation and Editing
- wavespeed-ai/qwen-image-2.0/text-to-image: Fast, high-quality text-to-image model with strong prompt fidelity, detailed rendering, and balanced performance for everyday creative tasks.
- wavespeed-ai/qwen-image-2.0-pro/text-to-image: Premium text-to-image model with enhanced detail, superior aesthetic quality, and finer control over complex multi-subject compositions.
- wavespeed-ai/qwen-image-2.0/edit: Intelligent image editing model for quick modifications, style adjustments, targeted content changes, and everyday creative workflows.
- wavespeed-ai/qwen-image-2.0-pro/edit: Advanced image editing model with higher precision, better context awareness, and production-grade output for professional retouching and creative transformation.
LoRA-ready Image Editing & Generation
- qwen-image/edit-plus-lora: Advanced image editing model with LoRA support, enabling precise style transfer, character customization, and high-fidelity local edits driven by text prompts.
- qwen-image/edit-lora: Lightweight edit model for LoRA-based style and character control, ideal for quick retouching, outfit changes, and consistent persona updates.
- qwen-image/text-to-image-lora: LoRA-enabled text-to-image generation that supports custom styles and characters while keeping strong prompt adherence and clean composition.
- jib-mix-qwen-image/text-to-image-lora: Mixed-style LoRA T2I model tuned for vivid anime and illustration aesthetics, combining sharp linework with rich color and expressive characters.
- qwen-image-lora-trainer: Training endpoint for building your own Qwen Image LoRA adapters from reference images, enabling personalized styles and characters across all LoRA-capable Qwen models.
Base Image Editing
- qwen-image/edit-plus: Enhanced image editing model for high-quality global and local edits, improving lighting, realism, and detail while preserving subject identity.
- qwen-image/edit: General-purpose edit model for everyday photo and artwork adjustments—ideal for quick fixes, background tweaks, and light retouching.
- qwen-image/edit-2511: High-consistency image editing model for reliable multi-subject, identity-preserving edits, delivering reduced drift, stronger geometric control, and cleaner, product-grade results for iterative, production workflows.
- qwen-image/edit-2511-edit-lora: LoRA-enhanced editing model built on the 2511 backbone—enables style injection, character customization, and fine-tuned aesthetic control while preserving the core stability of production-grade edits.
- qwen-image-max/edit: Advanced image editing model offering precise object manipulation, seamless background replacement, and intelligent style transfer, while preserving high-fidelity details and natural lighting.
Base Text-to-Image Generation
- qwen-image/text-to-image: Core T2I model that generates clean, realistic images from text prompts, suitable for product shots, portraits, and general creative use.
- jib-mix-qwen-image/text-to-image: Stylized T2I variant blending anime and illustration styles, producing vibrant, character-focused art with strong visual appeal.
- qwen-image/text-to-image-2512: Next-generation text-to-image model with enhanced prompt adherence, refined detail rendering, and improved compositional accuracy—engineered for photorealistic outputs and complex multi-element scene generation.
- qwen-image-max/text-to-image: Premium text-to-image model delivering exceptional detail, superior photorealism, and complex scene coherence. Designed for professional-grade generation with advanced lighting, texture rendering, and precise compositional control.
Utilities & Audio
- qwen-image/translate: Image translation utility that reads charts, UI screenshots, and text-heavy graphics, then outputs translated content while preserving layout semantics.
- qwen3-tts family: Fast text-to-speech model for natural-sounding voice previews, optimized for low latency in assistants, demos, and real-time applications.
































