Wan 2.1 Text to Image | High-Quality Text-to-Image API

Seedream 5.0 Pro уже здесь | Попробуйте в Генераторе изображений →

Панель управления Обзор AI-генераторHot Десктоп-приложение

LLM

Настройки

Главная/Обзор/WaveSpeed/Wan 2.1/Text To Image

wavespeed-ai /

Wan 2.1 Text-to-Image delivers ultra-realistic photographic images by adapting the Wan 2.1 video model for SOTA visual fidelity. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

text-to-image

Ввод

Enable Safety Checker

Ожидание

Envision an ethereal and highly decorative portrait of an androgynous Elven Monarch, seated upon a throne carved from living, iridescent wood within a moonlit glade. Their form is framed by the elegant, sinuous 'whiplash' curves characteristic of Art Nouveau. Long, flowing silver hair, intricately braided with glowing flora and pearls, cascades around them. They are adorned in gossamer robes of silk and moonlight, featuring delicate, repeating patterns of lilies and dragonfly wings. Their expression is serene and ancient, with eyes holding a gentle, knowing light. One hand elegantly gestures, causing magical, opalescent petals to swirl in the air. The background is a flat, decorative tapestry of intertwined vines, stylized trees, and celestial motifs, rendered in a soft palette of muted lavenders, sage greens, and creamy golds, with intricate gold leaf detailing. The entire composition is a harmonious symphony of organic forms and graceful lines, celebrating beauty, nature, and magic with the quintessential elegance of the Art Nouveau masters.

$0.02за запуск·~50 / $1

Envision an ethereal and highly decorative portrait of an androgynous Elven Monarch, seated upon a throne carved from living, iridescent wood within a moonlit glade. Their form is framed by the elegant, sinuous 'whiplash' curves characteristic of Art Nouveau. Long, flowing silver hair, intricately braided with glowing flora and pearls, cascades around them. They are adorned in gossamer robes of silk and moonlight, featuring delicate, repeating patterns of lilies and dragonfly wings. Their expression is serene and ancient, with eyes holding a gentle, knowing light. One hand elegantly gestures, causing magical, opalescent petals to swirl in the air. The background is a flat, decorative tapestry of intertwined vines, stylized trees, and celestial motifs, rendered in a soft palette of muted lavenders, sage greens, and creamy golds, with intricate gold leaf detailing. The entire composition is a harmonious symphony of organic forms and graceful lines, celebrating beauty, nature, and magic with the quintessential elegance of the Art Nouveau masters.

A candid, spontaneous selfie of a stunning young Black woman standing on her balcony at golden hour, dressed in a cropped white tee and wearing gold hoop earrings that catch the warm sunlight. Her curly hair frames her glowing skin naturally, illuminated by the soft last rays of the sun. The blurred city rooftops behind her fade gently into an orange-toned twilight sky. The image features authentic skin texture with subtle highlights and natural shadows, casual framing with a slight tilt capturing the intimate and effortless moment. The overall lighting and ambience reflect typical warm, natural light of an iPhone photo, making the scene feel genuine and elegantly powerful.

A platinum bob slips loose from a claw-clip, framing a woman’s face with frosted lips and a smudge of sunlit freckles. Her hand presses lightly to her cheek, stacked with chunky silver rings that glint like molten metal, bending in fluid, bold shapes. Black wrap-around sunglasses catch the faint reflection of city glass and the phone’s screen, while cool overcast light throws soft shadows on skin with pores and the fuzz of a wool rib cuff nearby. The shot tilts just off-center, fall-off blur wrapping around her cheek and shading, capturing the casual disarray of an artful moment cropped tight—close-up captured on Iphone, hand-face jewelry focus

A side-view photo of a cat walking gracefully along a narrow balcony railing at night. The background reveals a softly blurred city skyline glowing with lights—windows, streetlamps, and distant cars forming a bokeh effect. The cat's fur catches subtle reflections from the urban glow, and its tail balances high as it steps with precision. Cinematic night lighting, shallow depth of field, high-resolution photograph.

Intense medieval battle scene with knights clashing in close combat, swords swinging, and shields raised. The image is filled with dynamic motion blur — blurred swords, flying debris, and rushing figures convey the chaos of the fight. Dust rises from the ground, kicked up by charging horses and running soldiers. Armor glints in the sunlight, partially obscured by blur and dirt. The composition captures the raw energy of the battlefield, with blurred foreground action and slightly sharper figures in mid-ground. Gritty, cinematic lighting, overcast sky, and a muted, earth-toned color palette. Shot with a wide lens, slightly tilted, as if captured in the middle of battle.

Close-up of a woman’s hand fully submerged in clear ocean water, elegantly holding a bright yellow lemon. Her nails are clean and simple. Around the hand, small tropical fish swim gracefully, with strands of seaweed and soft marine plants drifting nearby. Sunlight filters through the water surface above, casting dappled light and gentle caustics on the skin and surrounding sea life. The skin has hyperrealistic detail — soft, luxurious, and radiant. The water is turquoise-green, slightly hazy with floating particles, creating a dreamy, cinematic underwater atmosphere. The scene feels elegant, epic, and editorial — a surreal blend of luxury and nature.

High‑quality photo. Black woman with a big afro leans on a mustard‑yellow 1970s convertible at a retro gas station. She wears high‑waisted orange flared pants and a tucked‑in paisley shirt that shows her waist. Large gold hoop earrings move slightly in the desert breeze. Warm late‑afternoon light falls on the glossy car and her face. Dusty asphalt, vintage pumps, and faded signs appear sharp. Pastel dusk sky fills the background. Shot eye‑level on a 50 mm lens, Kodak Gold 100 film with visible grain and a touch of lens flare. Late‑70s / early‑80s cinematic look.

A european woman with short, tousled hair leans in close to a man, her eyes gently closed. She wears a glowing red sweater; he has a jeans jacket with a bright collar. Golden ambient light softens their skin tones. Their faces are calm and close, framed tightly in an intimate, cinematic shot. The background is a gentle blur of city motion, adding contrast to their stillness. Subtle film grain evokes a timeless, romantic feel.

Close-up, top-down view of a teenage girl lying on a vibrant picnic blanket spread out on the green lawn of a sunny backyard garden. She’s laughing joyfully as a playful puppy stands beside her head, licking her ear. Her eyes are closed from laughter, and her expression is full of pure delight. The blanket is colorful — with bright patterns like florals or stripes — contrasting against the lush grass. Around them, soft natural sunlight filters through tree leaves, casting dappled shadows and warm highlights across her face, hair, and the puppy’s fur. Flowers, garden plants, or scattered toys add playful detail in the background. A lively, heartwarming outdoor moment filled with summer energy and color.

Close-up of a woman in ancient costume, with soft light falling on her skin, outlining delicate contours.

A close-up portrait of a cheerful Scandinavian man with bright blue eyes, enjoying an outdoor coffee in a quaint European town square bathed in warm morning light. The scene is bright and inviting, with crisp focus on his happy expression.

README

wan-2.1/text-to-image

Wan 2.1 is part of the Wan 2.1 foundation model suite, an advanced AI system developed to redefine video and image generation. This model focuses on text-to-image synthesis — transforming detailed written prompts into vivid, high-resolution visuals with cinematic precision.

🌟 Key Features

🎨 SOTA Image Quality Built on Wan 2.1’s next-generation video foundation, this model produces exceptional still-frame quality with realistic lighting, texture, and depth.
🧠 Multilingual Understanding Supports both Chinese and English prompts, ensuring accurate and context-rich image generation across languages.
⚙️ Fine Control with Parameters Adjustable inputs such as strength, width, and height provide creators with direct control over composition and style.
🪄 Powerful Visual Consistency Based on Wan-VAE, enabling coherent detail, color fidelity, and stylistic alignment across resolutions.
💰 Lightweight and Efficient High-quality generation at a base cost of just $0.02 per image, ideal for scalable creative workflows.

⚙️ Parameters

Parameter	Description
prompt*	Text description of the image to be generated (supports CN/EN).
image	(Optional) Upload a reference image for guided generation.
strength	Controls how strongly the image follows the prompt or reference (0–1).
size (width / height)	Define custom output resolution; max recommended ratio 2:1.
seed	Fix for reproducibility or randomize for variation.
output_format	Choose from `jpeg`, `png`, or `webp`.

💡 Example Prompt

Envision an ethereal and highly decorative portrait of an androgynous Elven Monarch, seated upon a throne carved from living iridescent wood within a moonlit glade. Intricate Art Nouveau details, luminous textures, soft-focus background, cinematic lighting.

💰 Pricing

Metric	Price
Per image generated	$0.02 / image

🎯 Use Cases

Concept Art & Illustration — Generate fantasy, sci-fi, or cinematic character art.
Visual Design & Branding — Create unique imagery for marketing, web, or product visuals.
Research & Visualization — Produce clear, detailed concept visuals from descriptive text.
Previsualization — Generate cinematic stills for film, animation, or game design workflows.

Примечание:Этот сайт использует модели ИИ, предоставляемые третьими лицами.

Wan 2.1 Text To Image API — Quick start

Grab a WaveSpeedAI API key, then call POST https://api.wavespeed.ai/api/v3/wavespeed-ai/wan-2.1/text-to-image with your input as JSON. The endpoint returns a prediction id. Start polling the result endpoint around every 2 seconds, increase the interval for long-running tasks, and stop on any terminal status. On completed, read output values from data.outputs. Examples for Wan 2.1 Text To Image below.

HTTP example

set -euo pipefail

: "${WAVESPEED_API_KEY:?Set WAVESPEED_API_KEY}"

REQUEST_BODY=$(cat <<'JSON'
{
    "prompt": "A cinematic shot of a city at sunset, soft golden light",
    "strength": 0.6,
    "size": "1024*1024",
    "seed": -1,
    "output_format": "jpeg"
}
JSON
)

# 1. Submit the prediction.
SUBMIT_RESPONSE=$(curl --silent --show-error --fail-with-body \
  -X POST "https://api.wavespeed.ai/api/v3/wavespeed-ai/wan-2.1/text-to-image" \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $WAVESPEED_API_KEY" \
  -d "$REQUEST_BODY")

TASK=$(printf '%s' "$SUBMIT_RESPONSE" | jq 'if has("data") then .data else . end')
PREDICTION_ID=$(printf '%s' "$TASK" | jq -r '.id')
if [ -z "$PREDICTION_ID" ] || [ "$PREDICTION_ID" = "null" ]; then
  printf 'Submission response did not contain a prediction id
' >&2
  exit 1
fi
RESULT_URL=$(printf '%s' "$TASK" | jq -r '.urls.get // empty')
if [ -z "$RESULT_URL" ]; then
  RESULT_URL="https://api.wavespeed.ai/api/v3/predictions/$PREDICTION_ID/result"
fi

# 2. Poll until the prediction finishes.
while true; do
  RESPONSE=$(curl --silent --show-error --fail-with-body "$RESULT_URL" \
    -H "Authorization: Bearer $WAVESPEED_API_KEY")
  RESULT=$(printf '%s' "$RESPONSE" | jq 'if has("data") then .data else . end')
  STATUS=$(printf '%s' "$RESULT" | jq -r '.status')
  case "$STATUS" in
    completed) printf '%s\n' "$RESULT" | jq '.outputs'; break ;;
    failed|cancelled|timeout) printf '%s\n' "$RESULT" | jq . >&2; exit 1 ;;
    created|processing) sleep 2 ;;
    *) printf 'Unexpected status: %s
' "$STATUS" >&2; exit 1 ;;
  esac
done

Node.js example

const submitUrl = "https://api.wavespeed.ai/api/v3/wavespeed-ai/wan-2.1/text-to-image";
const apiKey = process.env.WAVESPEED_API_KEY;
if (!apiKey) throw new Error('Set WAVESPEED_API_KEY');

async function requestJson(url, options = {}) {
  const response = await fetch(url, options);
  if (!response.ok) throw new Error(await response.text());
  return response.json();
}

// 1. Submit the prediction.
const body = await requestJson(submitUrl, {
  method: "POST",
  headers: {
    "Authorization": `Bearer ${apiKey}`,
    "Content-Type": "application/json",
  },
  body: JSON.stringify({
        "prompt": "A cinematic shot of a city at sunset, soft golden light",
        "strength": 0.6,
        "size": "1024*1024",
        "seed": -1,
        "output_format": "jpeg"
}),
});
const task = body.data ?? body;
if (!task.id) throw new Error("Submission response did not contain a prediction id");
const resultUrl = task.urls?.get ||
  `https://api.wavespeed.ai/api/v3/predictions/${task.id}/result`;

// 2. Poll until the prediction finishes.
while (true) {
  const resultBody = await requestJson(resultUrl, {
    headers: { "Authorization": `Bearer ${apiKey}` },
  });
  const result = resultBody.data ?? resultBody;
  if (result.status === "completed") {
    console.log(result.outputs);
    break;
  }
  if (["failed", "cancelled", "timeout"].includes(result.status)) throw new Error(JSON.stringify(result));
  if (!["created", "processing"].includes(result.status)) throw new Error("Unexpected status: " + result.status);
  await new Promise(resolve => setTimeout(resolve, 2000));
}

Python example

import json
import os
import time
from urllib.request import Request, urlopen

api_key = os.environ["WAVESPEED_API_KEY"]
headers = {"Authorization": f"Bearer {api_key}", "Content-Type": "application/json"}
payload = {
    "prompt": "A cinematic shot of a city at sunset, soft golden light",
    "strength": 0.6,
    "size": "1024*1024",
    "seed": -1,
    "output_format": "jpeg"
}

def request_json(url, data=None):
    request = Request(url, data=data, headers=headers, method="POST" if data else "GET")
    with urlopen(request) as response:
        return json.load(response)

# 1. Submit the prediction.
body = request_json("https://api.wavespeed.ai/api/v3/wavespeed-ai/wan-2.1/text-to-image", json.dumps(payload).encode())
task = body.get("data", body)
if not task.get("id"):
    raise RuntimeError("Submission response did not contain a prediction id")
result_url = task.get("urls", {}).get("get") or f"https://api.wavespeed.ai/api/v3/predictions/{task['id']}/result"

# 2. Poll until the prediction finishes.
while True:
    result_body = request_json(result_url)
    result = result_body.get("data", result_body)
    status = result.get("status")
    if status == "completed":
        print(result.get("outputs", []))
        break
    if status in {"failed", "cancelled", "timeout"}:
        raise RuntimeError(result)
    if status not in {"created", "processing"}:
        raise RuntimeError(f"Unexpected status: {status}")
    time.sleep(2)

Wan 2.1 Text To Image API — Frequently asked questions

What is the Wan 2.1 Text To Image API?

Wan 2.1 Text To Image is a WaveSpeedAI model for image generation, exposed as a REST API on WaveSpeedAI. Wan 2.1 Text-to-Image delivers ultra-realistic photographic images by adapting the Wan 2.1 video model for SOTA visual fidelity. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing. You can call it programmatically or try it from the playground above.

How do I call the Wan 2.1 Text To Image API?

POST your input parameters to the model's REST endpoint (shown in the API tab of this playground) with your WaveSpeedAI API key in the Authorization header. Submission returns a prediction ID. Poll the result endpoint starting around every 2 seconds, increase the interval for long-running tasks, and stop on any terminal status. The playground generates production-oriented Python, JavaScript, and cURL examples with timeouts, transient-error handling, and safe GET retries. Full request/response shape is documented at https://wavespeed.ai/docs/docs-api/wavespeed-ai/wan-2.1-text-to-image.

How much does Wan 2.1 Text To Image cost per run?

Wan 2.1 Text To Image starts at $0.020 per run. That figure is the base price — the final charge scales with the parameters you set in the form (output size, length, count, references, or whatever knobs this model exposes), so a higher-quality or larger output costs more than a minimal one. The exact cost for your current input is shown live next to the Generate button before you submit, and the actual per-call charge is recorded on the prediction afterwards.

What inputs does Wan 2.1 Text To Image accept?

Key inputs: `prompt`, `image`, `size`, `seed`, `enable_base64_output`, `enable_sync_mode`. The full JSON schema (types, defaults, allowed values) is rendered above the Generate button and mirrored in the API reference at https://wavespeed.ai/docs/docs-api/wavespeed-ai/wan-2.1-text-to-image.

How long does Wan 2.1 Text To Image take to generate?

Median end-to-end generation time on WaveSpeedAI is around 8 seconds per request, based on recent successful runs. Queue time varies with global demand; live status is visible in the prediction record.

Can I use Wan 2.1 Text To Image outputs commercially?

Commercial usage rights depend on the model's license, set by its provider (WaveSpeedAI). The license summary appears on the model card above; see WaveSpeedAI's Terms of Service for platform-level conditions.

ПримерыСмотреть всё

Похожие модели

README

wan-2.1/text-to-image

🌟 Key Features

⚙️ Parameters

💡 Example Prompt

💰 Pricing

🎯 Use Cases

Wan 2.1 Text To Image API — Quick start

Wan 2.1 Text To Image API — Frequently asked questions

Подробнее

Правовая информация

Ресурсы

Модели

Инструменты