Home/Explore/Wan 2.1 Video Models/wavespeed-ai/wan-2.1/text-to-image

text-to-image

wavespeed-ai/wan-2.1/text-to-image

Revolutionary text-to-image generation powered by Wan 2.1, delivering ultra-realistic images with photographic authenticity and exceptional detail fidelity. We modified the Wan 2.1 Video Model to make it also support image generation, and found it could achieve SOTA image generation quality!

Doc

Hint: You can drag and drop a file or click to upload

width
height
If enabled, the output will be encoded into a BASE64 string instead of a URL. This property is only available through the API.
If set to true, the safety checker will be enabled.
If set to true, the function will wait for the image to be generated and uploaded before returning the response. It allows you to get the image directly in the response. This property is only available through the API.

Idle

A platinum bob slips loose from a claw-clip, framing a woman’s face with frosted lips and a smudge of sunlit freckles. Her hand presses lightly to her cheek, stacked with chunky silver rings that glint like molten metal, bending in fluid, bold shapes. Black wrap-around sunglasses catch the faint reflection of city glass and the phone’s screen, while cool overcast light throws soft shadows on skin with pores and the fuzz of a wool rib cuff nearby. The shot tilts just off-center, fall-off blur wrapping around her cheek and shading, capturing the casual disarray of an artful moment cropped tight—close-up captured on Iphone, hand-face jewelry focus

Your request will cost $0.02 per run.

For $1 you can run this model approximately 50 times.

One more thing:

ExamplesView all

A platinum bob slips loose from a claw-clip, framing a woman’s face with frosted lips and a smudge of sunlit freckles. Her hand presses lightly to her cheek, stacked with chunky silver rings that glint like molten metal, bending in fluid, bold shapes. Black wrap-around sunglasses catch the faint reflection of city glass and the phone’s screen, while cool overcast light throws soft shadows on skin with pores and the fuzz of a wool rib cuff nearby. The shot tilts just off-center, fall-off blur wrapping around her cheek and shading, capturing the casual disarray of an artful moment cropped tight—close-up captured on Iphone, hand-face jewelry focus
A candid, spontaneous selfie of a stunning young Black woman standing on her balcony at golden hour, dressed in a cropped white tee and wearing gold hoop earrings that catch the warm sunlight. Her curly hair frames her glowing skin naturally, illuminated by the soft last rays of the sun. The blurred city rooftops behind her fade gently into an orange-toned twilight sky. The image features authentic skin texture with subtle highlights and natural shadows, casual framing with a slight tilt capturing the intimate and effortless moment. The overall lighting and ambience reflect typical warm, natural light of an iPhone photo, making the scene feel genuine and elegantly powerful.
A side-view photo of a cat walking gracefully along a narrow balcony railing at night. The background reveals a softly blurred city skyline glowing with lights—windows, streetlamps, and distant cars forming a bokeh effect. The cat's fur catches subtle reflections from the urban glow, and its tail balances high as it steps with precision. Cinematic night lighting, shallow depth of field, high-resolution photograph.
Intense medieval battle scene with knights clashing in close combat, swords swinging, and shields raised. The image is filled with dynamic motion blur — blurred swords, flying debris, and rushing figures convey the chaos of the fight. Dust rises from the ground, kicked up by charging horses and running soldiers. Armor glints in the sunlight, partially obscured by blur and dirt. The composition captures the raw energy of the battlefield, with blurred foreground action and slightly sharper figures in mid-ground. Gritty, cinematic lighting, overcast sky, and a muted, earth-toned color palette. Shot with a wide lens, slightly tilted, as if captured in the middle of battle.
Close-up of a woman’s hand fully submerged in clear ocean water, elegantly holding a bright yellow lemon. Her nails are clean and simple. Around the hand, small tropical fish swim gracefully, with strands of seaweed and soft marine plants drifting nearby. Sunlight filters through the water surface above, casting dappled light and gentle caustics on the skin and surrounding sea life. The skin has hyperrealistic detail — soft, luxurious, and radiant. The water is turquoise-green, slightly hazy with floating particles, creating a dreamy, cinematic underwater atmosphere. The scene feels elegant, epic, and editorial — a surreal blend of luxury and nature.
High‑quality photo. Black woman with a big afro leans on a mustard‑yellow 1970s convertible at a retro gas station. She wears high‑waisted orange flared pants and a tucked‑in paisley shirt that shows her waist. Large gold hoop earrings move slightly in the desert breeze. Warm late‑afternoon light falls on the glossy car and her face. Dusty asphalt, vintage pumps, and faded signs appear sharp. Pastel dusk sky fills the background. Shot eye‑level on a 50 mm lens, Kodak Gold 100 film with visible grain and a touch of lens flare. Late‑70s / early‑80s cinematic look.
A european woman with short, tousled hair leans in close to a man, her eyes gently closed. She wears a glowing red sweater; he has a jeans jacket with a bright collar. Golden ambient light softens their skin tones. Their faces are calm and close, framed tightly in an intimate, cinematic shot. The background is a gentle blur of city motion, adding contrast to their stillness. Subtle film grain evokes a timeless, romantic feel.
Close-up, top-down view of a teenage girl lying on a vibrant picnic blanket spread out on the green lawn of a sunny backyard garden. She’s laughing joyfully as a playful puppy stands beside her head, licking her ear. Her eyes are closed from laughter, and her expression is full of pure delight. The blanket is colorful — with bright patterns like florals or stripes — contrasting against the lush grass. Around them, soft natural sunlight filters through tree leaves, casting dappled shadows and warm highlights across her face, hair, and the puppy’s fur. Flowers, garden plants, or scattered toys add playful detail in the background. A lively, heartwarming outdoor moment filled with summer energy and color.
Envision an ethereal and highly decorative portrait of an androgynous Elven Monarch, seated upon a throne carved from living, iridescent wood within a moonlit glade. Their form is framed by the elegant, sinuous 'whiplash' curves characteristic of Art Nouveau. Long, flowing silver hair, intricately braided with glowing flora and pearls, cascades around them. They are adorned in gossamer robes of silk and moonlight, featuring delicate, repeating patterns of lilies and dragonfly wings. Their expression is serene and ancient, with eyes holding a gentle, knowing light. One hand elegantly gestures, causing magical, opalescent petals to swirl in the air. The background is a flat, decorative tapestry of intertwined vines, stylized trees, and celestial motifs, rendered in a soft palette of muted lavenders, sage greens, and creamy golds, with intricate gold leaf detailing. The entire composition is a harmonious symphony of organic forms and graceful lines, celebrating beauty, nature, and magic with the quintessential elegance of the Art Nouveau masters.
Close-up of a woman in ancient costume, with soft light falling on her skin, outlining delicate contours.
A close-up portrait of a cheerful Scandinavian man with bright blue eyes, enjoying an outdoor coffee in a quaint European town square bathed in warm morning light. The scene is bright and inviting, with crisp focus on his happy expression.

README

Wan 2.1 AI Video Model

We present Wan2.1, a comprehensive and open suite of video foundation models that pushes the boundaries of video generation. Wan2.1 offers these key features:

  • 👍 SOTA Performance: Wan2.1 consistently outperforms existing open-source models and state-of-the-art commercial solutions across multiple benchmarks.
  • 👍 Multiple Tasks: Wan2.1 excels in Text-to-Video, Image-to-Video, Video Editing, Text-to-Image, and Video-to-Audio, advancing the field of video generation.
  • 👍 Visual Text Generation: Wan2.1 is the first video model capable of generating both Chinese and English text, featuring robust text generation that enhances its practical applications.
  • 👍 Powerful Video VAE: Wan-VAE delivers exceptional efficiency and performance, encoding and decoding 1080P videos of any length while preserving temporal information, making it an ideal foundation for video and image generation.