Nano Banana 2.1 现已上线 — Google 最新模型 | 立即体验 →
Alibaba图片 API$0.02/次起

Qwen Image API

阿里巴巴 Qwen-Image——20B MMDiT 新一代文生图和编辑工具集,支持中英双语,具备多图编辑、LoRA 定制、分层合成和 96 种姿态的镜头角度系统。

文生图提供基础版、增强的 2512 和 2.0-pro 变体。编辑端点包括 Edit、Edit-Plus(多图、ControlNet)、Edit-LoRA、Edit-Multiple-Angles(96 姿态镜头系统)和 Layered(提示词引导的分解)。Qwen Image 2.0 系列变体也包含在同一前缀下。

Qwen Image 示例输出

概览

关于 Qwen Image API

Qwen Image 能做什么、它在 Alibaba 模型阵容中的定位,以及团队选择它的原因。

Qwen Image 是 Alibaba 推出的图像生成与编辑模型,可通过 WaveSpeedAI REST API 使用。阿里巴巴 Qwen-Image——20B MMDiT 新一代文生图和编辑工具集,支持中英双语,具备多图编辑、LoRA 定制、分层合成和 96 种姿态的镜头角度系统。

文生图提供基础版、增强的 2512 和 2.0-pro 变体。编辑端点包括 Edit、Edit-Plus(多图、ControlNet)、Edit-LoRA、Edit-Multiple-Angles(96 姿态镜头系统)和 Layered(提示词引导的分解)。Qwen Image 2.0 系列变体也包含在同一前缀下。

WaveSpeedAI 上的 Qwen Image 系列提供 21 个 REST 端点,涵盖 Image-To-Image, Text-To-Image, Training 个工作流。每个变体都有各自的定价、参数选项和示例输出——请选择与你的输入模态和生产约束相匹配的那一个,或使用同一个 API 密钥调用多个变体,组合成多步骤流水线。

使用与 WaveSpeedAI 上其他 1,000 多个 AI 模型相同的 API 密钥、账单账户和速率限制来运行 Qwen Image。无需单独对接供应商,无需各家 SDK,也无需应对各家不同的速率限制——一次集成即可覆盖从文生图、文生视频到音频合成、3D 生成、放大和编辑的全部能力。

端点

全部 Qwen Image API 端点

WaveSpeedAI 现已提供 21 个 Qwen Image 端点——请选择适合你工作流的变体。

image-to-imageQwen Image 2.0 Edit — Alibaba 的 Qwen Image image-to-image 预览

Qwen Image 2.0 Edit

Qwen Image 2.0 Edit is an advanced image-editing model with improved quality and better understanding of instructions. Up to 2k. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

$0.03 起
image-to-imageQwen Image 2.0 Pro Edit — Alibaba 的 Qwen Image image-to-image 预览

Qwen Image 2.0 Pro Edit

Qwen Image 2.0 Pro Edit is a professional-grade image editing model with superior quality and advanced instruction understanding. Up to 2k. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

$0.07 起
text-to-imageQwen Image 2.0 Text To Image — Alibaba 的 Qwen Image text-to-image 预览

Qwen Image 2.0 Text To Image

Qwen Image 2.0 is an advanced text-to-image model with enhanced image quality and improved prompt understanding. Up to 2k. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

$0.03 起
text-to-imageQwen Image 2.0 Pro Text To Image — Alibaba 的 Qwen Image text-to-image 预览

Qwen Image 2.0 Pro Text To Image

Qwen Image 2.0 Pro is a professional-grade text-to-image model with superior quality and advanced prompt understanding. Up to 2k. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

$0.07 起
image-to-imageQwen Image Edit 2509 Multiple Angles — Alibaba 的 Qwen Image image-to-image 预览

Qwen Image Edit 2509 Multiple Angles

Qwen Image Edit 2509 Multiple Angles is an AI image editing model that generates multiple-angle views of objects or scenes from a single image. Transform perspectives and create diverse viewpoints with text prompts. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

$0.025 起
image-to-imageQwen Image Max Edit — Alibaba 的 Qwen Image image-to-image 预览

Qwen Image Max Edit

Qwen Image Max Edit is an AI model for image editing with text prompts, supporting both Chinese and English languages. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

$0.07 起
text-to-imageQwen Image Max Text To Image — Alibaba 的 Qwen Image text-to-image 预览

Qwen Image Max Text To Image

Qwen Image Max is a text-to-image model with high-quality image generation supporting Chinese and English prompts. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

$0.07 起
image-to-imageQwen Image Edit Multiple Angles — Alibaba 的 Qwen Image image-to-image 预览

Qwen Image Edit Multiple Angles

Generate specific camera angles from a single image using a 96-pose camera system. Control horizontal rotation, vertical tilt, and zoom to create front, side, back views and more. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

$0.025 起
trainingQwen Image 2512 Lora Trainer — Alibaba 的 Qwen Image training 预览

Qwen Image 2512 Lora Trainer

Qwen-Image-2512 LoRA Trainer lets you train custom LoRA models 10x faster with style, character, and object training. From concept to model in minutes, not hours—upload a ZIP file containing images to start. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.

$1.00 起
text-to-imageQwen Image Text To Image 2512 Lora — Alibaba 的 Qwen Image text-to-image 预览

Qwen Image Text To Image 2512 Lora

Qwen-Image-2512 LoRA is an enhanced 20B MMDiT text-to-image model with LoRA support for fast customization and refined image generation. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.

$0.025 起
text-to-imageQwen Image Text To Image 2512 — Alibaba 的 Qwen Image text-to-image 预览

Qwen Image Text To Image 2512

Qwen Image 2512 is Qwen's latest text-to-image model with enhanced prompt understanding, superior text rendering, and versatile aspect ratio support. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.

$0.02 起
image-to-imageQwen Image Edit 2511 Lora — Alibaba 的 Qwen Image image-to-image 预览

Qwen Image Edit 2511 Lora

Qwen Image Edit 2511 LoRA is an enhanced version with custom LoRA support for personalized styles. It delivers stronger edit consistency, robust multi-person identity/pose consistency, custom LoRA styles, enhanced industrial/product design, and improved geometric reasoning for structure-preserving edits. Built for stable production use with a ready-to-use REST API, no cold starts, and predictable pricing.

$0.025 起
image-to-imageQwen Image Edit 2511 — Alibaba 的 Qwen Image image-to-image 预览

Qwen Image Edit 2511

Qwen Image Edit 2511 is a major upgrade over 2509 for real-world image editing and design. It delivers stronger edit consistency, robust multi-person identity/pose consistency, built-in LoRA styles, enhanced industrial/product design, and improved geometric reasoning for structure-preserving edits. Built for stable production use with a ready-to-use REST API, no cold starts, and predictable pricing.

$0.02 起
image-to-imageQwen Image Layered — Alibaba 的 Qwen Image image-to-image 预览

Qwen Image Layered

Qwen-Image Layered is a unified image-layer decomposition model for prompt-guided compositing. Provide points, boxes, or rough masks to isolate subjects and regions, and the model splits a single image into multiple RGBA layers with clean alpha, soft edges, and correct occlusion order. Ready-to-use REST inference API with fast response, no cold starts, and affordable pricing.

$0.025 起
image-to-imageQwen Image Edit Plus Lora — Alibaba 的 Qwen Image image-to-image 预览

Qwen Image Edit Plus Lora

Qwen-Image-Edit-Plus (2509) is 20B MMDiT image-to-image editor supporting multi-image edits, single-image consistency, and native ControlNet. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

$0.025 起
image-to-imageQwen Image Edit Plus — Alibaba 的 Qwen Image image-to-image 预览

Qwen Image Edit Plus

Qwen-Image-Edit-Plus (2509) is a 20B MMDiT image editor with multi-image editing, single-image consistency and native ControlNet support. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

$0.02 起
image-to-imageQwen Image Edit Lora — Alibaba 的 Qwen Image image-to-image 预览

Qwen Image Edit Lora

Qwen-Image-Edit LoRA (20B) enables bilingual Chinese/English image-to-image editing with style preservation and semantic and appearance edits. Ready-to-use REST API, best performance, no coldstarts, affordable pricing.

$0.025 起
image-to-imageQwen Image Edit — Alibaba 的 Qwen Image image-to-image 预览

Qwen Image Edit

Qwen-Image-Edit is a 20B MMDiT image-to-image model offering precise bilingual (Chinese & English) text edits while preserving style. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

$0.02 起
trainingQwen Image Lora Trainer — Alibaba 的 Qwen Image training 预览

Qwen Image Lora Trainer

Train custom Qwen-Image LoRA models 10x faster. Style training, character training, object training. From concept to model in minutes, not hours. Upload a ZIP file containing images to start!

$1.00 起
text-to-imageQwen Image Text To Image Lora — Alibaba 的 Qwen Image text-to-image 预览

Qwen Image Text To Image Lora

Qwen-Image LoRA is a 20B MMDiT next-gen text-to-image model with LoRA support for fast customization and refined image generation. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

$0.025 起
text-to-imageQwen Image Text To Image — Alibaba 的 Qwen Image text-to-image 预览

Qwen Image Text To Image

Qwen-Image is a 20B MMDiT next-gen text-to-image model that generates images from text prompts. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

$0.02 起

示例

看看 Qwen Image 的实际效果

由 Qwen Image API 生成的真实输出。悬停在任意视频上即可预览,点击可打开全尺寸查看器。

使用方法

如何使用 Qwen Image API

从注册到完成一次生成,只需四步。完整的 Python、Node.js 和 cURL 示例见下方的 API 部分。

  1. 01

    获取 API 密钥

    注册 WaveSpeedAI 账号,并从控制台复制你的 API 密钥。新账号附带免费体验额度——足够在开始计费前把 Playground 运行几十次。

  2. 02

    提交预测

    把你的输入以 JSON 形式 POST 到 https://api.wavespeed.ai/api/v3/wavespeed-ai/qwen-image/text-to-image。端点会立即返回预测 ID——生成是异步的,因此推理期间你无需保持连接。

  3. 03

    轮询完成状态

    GET https://api.wavespeed.ai/api/v3/predictions/{request_id}/result。状态为 completed 时返回输出;状态为 failed、cancelled、timeout 或 deleted 时以错误终止;其他任何状态都继续轮询。

  4. 04

    读取输出 URL

    状态变为 "completed" 后,从 data.outputs[0] 读取 URL。该 URL 指向 WaveSpeedAI CDN 上你生成的媒体——具体是图片、视频、音频还是 3D 文件,取决于你调用的 Qwen Image 变体。

应用场景

用 Qwen Image 可以构建什么

开发者和创作者使用 Qwen Image API 的常见工作流。

01

20B MMDiT 文生图

wavespeed-ai/qwen-image/text-to-image 是 20B MMDiT 新一代文生图模型,可根据文本提示词生成图像——是 Qwen Image 系列的基础生成端点。

text-to-image20bmmdit
02

文字渲染出众的增强版 2512

wavespeed-ai/qwen-image/text-to-image-2512 是 Qwen 最新的文生图模型,据目录介绍,提示词理解增强、文字渲染出众,并支持多种画幅比例。

2512typographyaspect-ratio
03

用 Edit-Plus 进行多图编辑

wavespeed-ai/qwen-image/edit-plus 是 20B MMDiT 编辑器,支持多图编辑、单图一致性和原生 ControlNet——适合需要参考多张源图像的复杂编辑。

edit-plusmulti-imagecontrolnet
04

96 种姿态的镜头角度控制

wavespeed-ai/qwen-image/edit-multiple-angles 利用 96 种姿态的镜头系统,从单张图像生成特定的镜头角度——控制水平旋转、垂直倾斜和缩放,得到正面、侧面、背面等视图。

camera-angles96-pose3d-views
05

分层合成分解

wavespeed-ai/qwen-image/layered 是统一的图像分层分解模型,用于提示词引导的合成——提供点、框或粗略蒙版来分离主体,把单张图像拆分为多个图层。

layeredcompositingdecomposition
06

LoRA 定制

支持 LoRA 的变体(text-to-image-2512-lora、edit-lora、edit-plus-lora)可实现快速定制和更精细的生成——训练或应用 LoRA 检查点,以保持风格、角色或品牌的一致性。

loracustomizationconsistency

技巧

Qwen Image 提示词技巧

让 Qwen Image 输出更好结果的实用建议——总结自生产流水线中各类图像模型都适用的做法。

  1. 01

    用 2512 获得最新的文字渲染

    text-to-image-2512 是增强变体,文字渲染和提示词理解更出色——排版重要时,请选它而不是基础的 text-to-image。

  2. 02

    用 Edit-Multiple-Angles 获得产品多角度视图

    edit-multiple-angles 可根据单张产品照片生成正面、侧面、背面和自定义的镜头角度——适合电商目录,无需多机位拍摄。

  3. 03

    多图参考使用 Edit-Plus

    当编辑依赖多张源图像时,请使用支持原生 ControlNet 的 edit-plus,而不是单图编辑。

  4. 04

    合成工作流使用 Layered

    使用 layered 把图像分解为提示词引导的图层——用点、框或粗略蒙版分离主体,用于下游合成。

  5. 05

    用 LoRA 变体保持品牌一致

    通过 text-to-image-2512-lora 或 edit-plus-lora 应用训练好的 LoRA 检查点,在多次生成中保持固定的风格、角色或品牌身份。

  6. 06

    中英双语编辑

    Edit 和 Edit-LoRA 支持中英双语的图生图编辑并保留风格——适合本地化工作流。

定价

Qwen Image API 定价

按输出计费。最终费用会随你在各变体 Playground 中设置的参数(分辨率、时长、输出数量、参考素材)而变化。

API

调用 Qwen Image API

在 wavespeed.ai/accesskey 注册并获取 API 密钥,然后通过 REST 提交预测。Playground 可以为任意输入组合生成可直接粘贴的示例代码。

POSThttps://api.wavespeed.ai/api/v3/wavespeed-ai/qwen-image/text-to-image

# 1. Submit the prediction.
SUBMIT_RESPONSE=$(curl --silent --show-error --fail-with-body \
  -X POST "https://api.wavespeed.ai/api/v3/wavespeed-ai/qwen-image/text-to-image" \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $WAVESPEED_API_KEY" \
  -d '{
    "prompt": "A cinematic shot of a city at sunset, soft golden light",
    "size": "1024*1024",
    "output_format": "jpeg"
}')

TASK=$(printf '%s' "$SUBMIT_RESPONSE" | jq 'if has("data") then .data else . end')
PREDICTION_ID=$(printf '%s' "$TASK" | jq -r '.id')
if [ -z "$PREDICTION_ID" ] || [ "$PREDICTION_ID" = "null" ]; then
  printf 'Submission response did not contain a prediction id
' >&2
  exit 1
fi
RESULT_URL="https://api.wavespeed.ai/api/v3/predictions/$PREDICTION_ID/result"

# 2. Poll until the prediction finishes.
while true; do
  RESPONSE=$(curl --silent --show-error --fail-with-body "$RESULT_URL" \
    -H "Authorization: Bearer $WAVESPEED_API_KEY")
  RESULT=$(printf '%s' "$RESPONSE" | jq 'if has("data") then .data else . end')
  STATUS=$(printf '%s' "$RESULT" | jq -r '.status')
  case "$STATUS" in
    completed) printf '%s\n' "$RESULT" | jq '.outputs'; break ;;
    failed|cancelled|timeout|deleted) printf '%s\n' "$RESULT" | jq . >&2; exit 1 ;;
    *) sleep 2 ;;
  esac
done

对比

Qwen Image 与其他方案对比

在 WaveSpeedAI 上,何时应选择 Qwen Image 而不是同类模型。

Qwen Image 对比 Seedream 4.5

Seedream 4.5 强调排版,并提供用于多图身份锁定的 Sequential 变体。Qwen Image 的编辑范围更广——Edit-Plus、多角度、分层合成,以及中英双语编辑。

Qwen Image 对比 GPT Image 2

GPT Image 2 有明确的质量档位和参考图像编辑工作流。Qwen Image 提供原生 ControlNet、96 种姿态的镜头角度和分层分解——编辑方式不同,单次调用成本更低。

Qwen Image 对比 Nano Banana 2

Nano Banana 2 提供多角色一致性(最多 5 个)和网页搜索增强。Qwen Image 的优势在于编辑深度——多图 Edit-Plus、镜头角度生成和提示词引导的图层分解。

常见问题

Qwen Image API — 常见问题

定价、许可、集成——关于在 WaveSpeedAI 上运行 Qwen Image 的常见问题。

Qwen Image API 是什么?

Qwen Image 是 Alibaba 的图像生成模型,在 WaveSpeedAI 上以 REST API 形式提供。阿里巴巴 Qwen-Image——20B MMDiT 新一代文生图和编辑工具集,支持中英双语,具备多图编辑、LoRA 定制、分层合成和 96 种姿态的镜头角度系统。你可以通过编程方式调用它,也可以在上方链接的 Playground 中试用。

如何调用 Qwen Image API?

注册 WaveSpeedAI 账号,从 /accesskey 复制你的 API 密钥,然后把输入以 JSON 形式 POST 到 https://api.wavespeed.ai/api/v3/wavespeed-ai/qwen-image/text-to-image。端点会返回预测 ID。从大约每 2 秒一次开始轮询结果端点,长耗时任务可适当拉长间隔,并在任何终止状态时停止。上方有面向生产环境的 Python / Node.js / cURL 示例。

Qwen Image API 的费用是多少?

Qwen Image 每次调用 $0.02 起。实际费用会随你设置的参数(分辨率、时长、输出数量、参考素材)而变化。Playground 中“生成”按钮旁的实时费用预览会显示你当前输入对应的准确价格。

有哪些 Qwen Image 变体可用?

WaveSpeedAI 托管了 21 个已上线的 Qwen Image 端点:wavespeed-ai/qwen-image-2.0/edit, wavespeed-ai/qwen-image-2.0-pro/edit, wavespeed-ai/qwen-image-2.0/text-to-image, wavespeed-ai/qwen-image-2.0-pro/text-to-image, wavespeed-ai/qwen-image/edit-2509-multiple-angles, wavespeed-ai/qwen-image-max/edit, wavespeed-ai/qwen-image-max/text-to-image, wavespeed-ai/qwen-image/edit-multiple-angles等。每个变体都有自己的 Playground 页面和定价。

Qwen Image 的输出可以商用吗?

商用权利遵循 Alibaba 的模型许可。大多数 Alibaba 模型允许商用输出;具体许可摘要请查看各模型的 Playground 页面,平台层面的条件请参阅 WaveSpeedAI 的服务条款。

为什么要在 WaveSpeedAI 上使用 Qwen Image,而不是直接对接?

一个 API 密钥、一个账单账户,即可使用 Qwen Image 以及来自其他提供方的 1,000 多个 AI 模型。无需逐家配置 SDK,无需应对各自独立的速率限制,也无需为每家重写集成代码。价格通常与 Alibaba 直接提供的 API 持平或更低。

提供方

关于 Alibaba

Qwen Image 及 WaveSpeedAI 上 Alibaba 更多模型背后的团队。

阿里巴巴通义实验室推出了 Wan 系列视频模型和 Qwen 系列大语言模型。Wan 以开放权重发布、变体覆盖广泛(文生视频、图生视频、参考生视频、视频编辑、视频延长、图像编辑、文生图)而著称,并在运动稳定性和跨多语言提示词的提示词遵循方面表现稳定。

在 WaveSpeedAI 上用 Qwen Image 开始构建

注册即送免费体验额度。一个 API 密钥,即可使用来自 Alibaba 及其他所有提供方的 1,000 多个 AI 模型。