Seedance 2.5 Now Live | Try in Video Generator →
Home/Explore/Character Ai/Ovi/Image To Video

character-ai/

Ovi is a Veo-3-like image-to-video model that generates synchronized video and audio from text or text+image prompts. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

image-to-video
Input

Idle

$0.15per run·~66 / $10

Next:

ExamplesView all

Bright evenly lit laboratory room with metallic walls and soft white light reflections. A human man in a suit stands face-to-face with a humanoid robot, both in perfect focus. Camera: static medium close-up, centered framing, high exposure with clear details on both faces. Mood: tense, thoughtful, futuristic. <S>We built you to understand us.<E> A Sign <S>But sometimes I wonder if you understand us too well.<E> The robot tilts its head slightly, eyes glowing faint blue, voice calm and precise. <S>Understanding is not the same as becoming.<E> <AUDCAP>Soft ambient hum of electronics, faint mechanical servo sounds, two clear voices — human and synthetic, calm and steady<ENDAUDCAP>

A 5-second, dynamic close-up of a sleek, advanced android's head and upper torso. Its armored plates are etched with neon circuit patterns that pulse with a soft blue light. Its face is a polished metal and dark glass visor. As it boots up, its articulated jaw and vocal synthesizer move with precise, mechanical motion to form the words. Mood: Technological, mysterious, and immersive. <S>System. Online.<E> <AUDCAP>The clear, synthetic voice of the android, the low hum of its internal systems, and the faint, distant sound of hovering vehicles and city rain.<ENDAUDCAP>

A 5-second, static shot of a kind old Ghibli-style man with a wrinkled face and gentle eyes. He is seated at his workbench, holding a small wooden toy. He looks up and speaks softly to the viewer, his mouth moving clearly to form the words. The style is soft watercolor and pastel. Mood: Peaceful, wise, and nostalgic. <S>Just a little more...<E> <AUDCAP>The soft, raspy voice of the old man, the gentle sound of a breeze, and the distant chime of a wind bell.<ENDAUDCAP>

Raised his hand and said hello

A bearded man wearing large dark sunglasses and a blue patterned cardigan sits in a studio, actively speaking into a large, suspended microphone. He has headphones on and gestures with his hands, displaying rings on his fingers. Behind him, a wall is covered with red, textured sound-dampening foam on the left, and a white banner on the right features the "CHOICE FM" logo and various social media handles like "@ilovechoicefm" with "RALEIGH" below it. The man intently addresses the microphone, articulating, <S>is talent. It's all about authenticity. You gotta be who you really are, especially if you're working<E>. He leans forward slightly as he speaks, maintaining a serious expression behind his sunglasses.. <AUDCAP>Clear male voice speaking into a microphone, a low background hum.<ENDAUDCAP>

A medium close-up of a young woman standing on a sun-drenched hilltop at golden hour. A gentle breeze blows through her hair. She is turning towards the camera with a radiant, genuine smile, her mouth perfectly formed in the middle of the word "beautiful". Her eyes are squinting slightly against the low sun, filled with contentment. Camera: Static shot, sharp focus on her face and mouth. Bright, natural daylight, high exposure with soft shadows that define facial features. Mood: Serene, joyful, cinematic realism. <S>What a beautiful day!<E> <AUDCAP>The gentle rustle of leaves in the wind, distant chirping of birds, her clear and happy voice, a faint sigh of contentment after speaking.<ENDAUDCAP>

Related Models

README

Ovi (I2V Version)

Ovi is a veo-3 like, image-to-audio-video (I2AV) generation model that creates synchronized video and audio from a single image plus a descriptive text prompt.

It is designed for short-form storytelling, where a still image is brought to life with cinematic motion, dialogue, and sound.

🌟 Key Features

  • 🎬 Image → Video+Audio – Bring a static image to life with synchronized audiovisual output.
  • 📝 Prompt-driven – Use text prompts to control scene dynamics, style, and audio.
  • 🗣️ Speech & Sound – Insert dialogue or sound effects using special tags.
  • ⏱️ Short-form Output – Generates 5-second clips at 24 FPS.

💲 Pricing

Video LengthCost
5 seconds$0.15

Billing Rules

  • Minimum charge: 5 seconds

🎨 How to Use

  1. Upload Image
  • Provide a reference image as the base frame.
  • Make sure the URL is valid and accessible (a preview should appear).
  1. Enter Prompt
  • Describe scene motion, style, and atmosphere.

  • Use tags for sound:

  • <S>... <E> → Speech (converted into spoken audio)

  • <AUDCAP>... <ENDAUDCAP> → Background audio / effects

  1. Set Seed
  • -1 = random output
  • Any fixed number = reproducible results
  1. Run
  • Click Run $0.15 to generate your 5s image-to-audio-video clip.
  • Preview and download the result.

📝 Prompt Example

A wide shot of a medieval knight standing in the rain, sword planted into the ground, glowing with mystical energy. 
<S>I will defend this land until my last breath.<E> 
<AUDCAP>Thunder rolls across the dark sky, distant war drums echo.<ENDAUDCAP>

🙏 Acknowledgements

  • Wan2.2 – Video backbone initialization
  • MMAudio – Audio encoder/decoder inspiration

⭐ Citation

If Ovi is useful, please ⭐ the repo and cite the paper:

@misc{low2025ovitwinbackbonecrossmodal,
 title={Ovi: Twin Backbone Cross-Modal Fusion for Audio-Video Generation}, 
 author={Chetwin Low and Weimin Wang and Calder Katyal},
 year={2025},
 eprint={2510.01284},
 archivePrefix={arXiv},
 primaryClass={cs.MM},
 url={https://arxiv.org/abs/2510.01284}, 
}
Note:This website uses AI models provided by third parties. Documentation prices are for reference and may be outdated. The Generate button shows an estimate; the final task charge prevails.

Ovi Image To Video API — Quick start

Grab a WaveSpeedAI API key, then call POST https://api.wavespeed.ai/api/v3/character-ai/ovi/image-to-video with your input as JSON. The endpoint returns a prediction id. Start polling the result endpoint around every 2 seconds, increase the interval for long-running tasks, and stop on any terminal status. On completed, read output values from data.outputs. Examples for Ovi Image To Video below.

HTTP example
set -euo pipefail

: "${WAVESPEED_API_KEY:?Set WAVESPEED_API_KEY}"

REQUEST_BODY=$(cat <<'JSON'
{
    "prompt": "A cinematic shot of a city at sunset, soft golden light",
    "image": "https://interactive-examples.mdn.mozilla.net/media/cc0-images/painted-hand-298-332.jpg",
    "seed": -1
}
JSON
)

# 1. Submit the prediction.
SUBMIT_RESPONSE=$(curl --silent --show-error --fail-with-body \
  -X POST "https://api.wavespeed.ai/api/v3/character-ai/ovi/image-to-video" \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $WAVESPEED_API_KEY" \
  -d "$REQUEST_BODY")

TASK=$(printf '%s' "$SUBMIT_RESPONSE" | jq 'if has("data") then .data else . end')
PREDICTION_ID=$(printf '%s' "$TASK" | jq -r '.id')
if [ -z "$PREDICTION_ID" ] || [ "$PREDICTION_ID" = "null" ]; then
  printf 'Submission response did not contain a prediction id
' >&2
  exit 1
fi
RESULT_URL=$(printf '%s' "$TASK" | jq -r '.urls.get // empty')
if [ -z "$RESULT_URL" ]; then
  RESULT_URL="https://api.wavespeed.ai/api/v3/predictions/$PREDICTION_ID/result"
fi

# 2. Poll until the prediction finishes.
while true; do
  RESPONSE=$(curl --silent --show-error --fail-with-body "$RESULT_URL" \
    -H "Authorization: Bearer $WAVESPEED_API_KEY")
  RESULT=$(printf '%s' "$RESPONSE" | jq 'if has("data") then .data else . end')
  STATUS=$(printf '%s' "$RESULT" | jq -r '.status')
  case "$STATUS" in
    completed) printf '%s\n' "$RESULT" | jq '.outputs'; break ;;
    failed|cancelled|timeout) printf '%s\n' "$RESULT" | jq . >&2; exit 1 ;;
    created|processing) sleep 2 ;;
    *) printf 'Unexpected status: %s
' "$STATUS" >&2; exit 1 ;;
  esac
done
Node.js example
const submitUrl = "https://api.wavespeed.ai/api/v3/character-ai/ovi/image-to-video";
const apiKey = process.env.WAVESPEED_API_KEY;
if (!apiKey) throw new Error('Set WAVESPEED_API_KEY');

async function requestJson(url, options = {}) {
  const response = await fetch(url, options);
  if (!response.ok) throw new Error(await response.text());
  return response.json();
}

// 1. Submit the prediction.
const body = await requestJson(submitUrl, {
  method: "POST",
  headers: {
    "Authorization": `Bearer ${apiKey}`,
    "Content-Type": "application/json",
  },
  body: JSON.stringify({
        "prompt": "A cinematic shot of a city at sunset, soft golden light",
        "image": "https://interactive-examples.mdn.mozilla.net/media/cc0-images/painted-hand-298-332.jpg",
        "seed": -1
}),
});
const task = body.data ?? body;
if (!task.id) throw new Error("Submission response did not contain a prediction id");
const resultUrl = task.urls?.get ||
  `https://api.wavespeed.ai/api/v3/predictions/${task.id}/result`;

// 2. Poll until the prediction finishes.
while (true) {
  const resultBody = await requestJson(resultUrl, {
    headers: { "Authorization": `Bearer ${apiKey}` },
  });
  const result = resultBody.data ?? resultBody;
  if (result.status === "completed") {
    console.log(result.outputs);
    break;
  }
  if (["failed", "cancelled", "timeout"].includes(result.status)) throw new Error(JSON.stringify(result));
  if (!["created", "processing"].includes(result.status)) throw new Error("Unexpected status: " + result.status);
  await new Promise(resolve => setTimeout(resolve, 2000));
}
Python example
import json
import os
import time
from urllib.request import Request, urlopen

api_key = os.environ["WAVESPEED_API_KEY"]
headers = {"Authorization": f"Bearer {api_key}", "Content-Type": "application/json"}
payload = {
    "prompt": "A cinematic shot of a city at sunset, soft golden light",
    "image": "https://interactive-examples.mdn.mozilla.net/media/cc0-images/painted-hand-298-332.jpg",
    "seed": -1
}

def request_json(url, data=None):
    request = Request(url, data=data, headers=headers, method="POST" if data else "GET")
    with urlopen(request) as response:
        return json.load(response)

# 1. Submit the prediction.
body = request_json("https://api.wavespeed.ai/api/v3/character-ai/ovi/image-to-video", json.dumps(payload).encode())
task = body.get("data", body)
if not task.get("id"):
    raise RuntimeError("Submission response did not contain a prediction id")
result_url = task.get("urls", {}).get("get") or f"https://api.wavespeed.ai/api/v3/predictions/{task['id']}/result"

# 2. Poll until the prediction finishes.
while True:
    result_body = request_json(result_url)
    result = result_body.get("data", result_body)
    status = result.get("status")
    if status == "completed":
        print(result.get("outputs", []))
        break
    if status in {"failed", "cancelled", "timeout"}:
        raise RuntimeError(result)
    if status not in {"created", "processing"}:
        raise RuntimeError(f"Unexpected status: {status}")
    time.sleep(2)

Ovi Image To Video API — Frequently asked questions

What is the Ovi Image To Video API?

Ovi Image To Video is a Character Ai model for video generation from images, exposed as a REST API on WaveSpeedAI. Ovi is a Veo-3-like image-to-video model that generates synchronized video and audio from text or text+image prompts. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing. You can call it programmatically or try it from the playground above.

How do I call the Ovi Image To Video API?

POST your input parameters to the model's REST endpoint (shown in the API tab of this playground) with your WaveSpeedAI API key in the Authorization header. Submission returns a prediction ID. Poll the result endpoint starting around every 2 seconds, increase the interval for long-running tasks, and stop on any terminal status. The playground generates production-oriented Python, JavaScript, and cURL examples with timeouts, transient-error handling, and safe GET retries. Full request/response shape is documented at https://wavespeed.ai/docs/docs-api/character-ai/character-ai-ovi-image-to-video.

How much does Ovi Image To Video cost per run?

Ovi Image To Video starts at $0.15 per run. That figure is the base price — the final charge scales with the parameters you set in the form (output size, length, count, references, or whatever knobs this model exposes), so a higher-quality or larger output costs more than a minimal one. The exact cost for your current input is shown live next to the Generate button before you submit, and the actual per-call charge is recorded on the prediction afterwards.

What inputs does Ovi Image To Video accept?

Key inputs: `prompt`, `image`, `seed`. The full JSON schema (types, defaults, allowed values) is rendered above the Generate button and mirrored in the API reference at https://wavespeed.ai/docs/docs-api/character-ai/character-ai-ovi-image-to-video.

How long does Ovi Image To Video take to generate?

Median end-to-end generation time on WaveSpeedAI is around 115 seconds per request, based on recent successful runs. Queue time varies with global demand; live status is visible in the prediction record.

Can I use Ovi Image To Video outputs commercially?

Commercial usage rights depend on the model's license, set by its provider (Character Ai). The license summary appears on the model card above; see WaveSpeedAI's Terms of Service for platform-level conditions.

Ovi Image to Video | Fast Image-to-Video API on WaveSpeedAI