Google Gemini Omni 1.1 Flash Image To Video API Documentation

Google Gemini Omni 1.1 Flash Image To Video API Documentation

Playground

Try it on WaveSpeedAI!

Gemini Omni 1.1 Flash Image-to-Video animates a start image into short AI videos, optionally targeting an end frame and generating synchronized audio at resolutions from 360P to 4K for social content, creative storytelling, marketing videos, and production workflows. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

Features

Gemini Omni 1.1 Flash Image-to-Video animates a start image into a video with synchronized audio. You can optionally add a final frame to guide the ending of the shot, making it useful for controlled transitions, reveal sequences, camera movement, and other prompt-guided image-to-video workflows.


Why Choose This?

  • Start-frame image animation
    Turn a single image into a video with motion, scene progression, and synchronized audio.

  • Optional end-frame guidance
    Add last_image when you want more control over how the shot ends.

  • Smooth transition generation
    Useful for camera moves, reveal shots, visual transformations, and controlled motion between two keyframes.

  • Draft-to-production workflow
    Use 360p for quick testing, then move to 1080p or 4k for higher-quality output.

  • Synchronized audio
    Generate motion and matching audio together in one workflow.

  • Flexible output formats
    Supports both 16:9 and 9:16 for landscape and vertical video creation.


Parameters

ParameterRequiredDescription
imageYesStart-frame image URL.
last_imageNoOptional end-frame image URL used to guide the ending of the video.
promptYesMotion, camera movement, transition, scene, style, and audio instructions.
aspect_ratioNoOutput aspect ratio: 16:9 or 9:16. Default: 16:9.
resolutionNoOutput resolution: 360p, 720p, 1080p, or 4k. Default: 720p.
durationNoVideo duration in seconds. Range: 3–10. Default: 8.

How to Use

  1. Upload the start image — Provide the image that should define the first frame.
  2. Add an end image optional — Use last_image when you want to guide the final frame or overall transition direction.
  3. Write your prompt — Describe the motion, camera movement, transition behavior, scene changes, style, and audio.
  4. Choose aspect ratio — Use 16:9 for landscape video or 9:16 for vertical content.
  5. Select resolution — Use 360p for drafts, 720p for standard output, or 1080p / 4k for higher-quality results.
  6. Set duration — Choose a duration from 3 to 10 seconds.
  7. Submit — Generate the final video with synchronized audio.

Pricing

Pricing is based on generated video duration and selected resolution.

ResolutionPrice per second
360p$0.03
720p$0.10
1080p$0.15
4k$0.30

Example Costs

Resolution5s8s10s
360p$0.15$0.24$0.30
720p$0.50$0.80$1.00
1080p$0.75$1.20$1.50
4k$1.50$2.40$3.00

Best Use Cases

  • Image animation — Turn still images into short motion clips with audio.
  • Controlled transitions — Use image and last_image to guide the beginning and end of a shot.
  • Reveal and transformation scenes — Create visual change between two defined states.
  • Social media content — Generate vertical or landscape motion content from static artwork or product images.
  • Creative prototyping — Quickly test motion, pacing, and shot progression from a single keyframe.
  • Marketing and campaign assets — Animate posters, product visuals, and branded images into short videos.

Pro Tips

  • Use a clear, high-quality start image for better subject consistency.
  • Add last_image when the final pose, composition, or scene state matters.
  • Describe both what moves and how it moves in the prompt.
  • Include camera instructions such as push-in, orbit, pan, or pull-back for more controlled results.
  • Mention ambience, music, dialogue, or sound effects directly when audio guidance matters.
  • Start with 360p for fast iteration, then move to 1080p or 4k for final output.
  • Use shorter durations for quick tests and longer durations when the transition needs more time to develop.

Authentication

For authentication details, please refer to the Authentication Guide.

API Endpoints

Submit Task & Query Result

set -euo pipefail

export WAVESPEED_API_KEY="your-api-key"

REQUEST_BODY=$(cat <<'JSON'
{
  "image": "https://interactive-examples.mdn.mozilla.net/media/cc0-images/painted-hand-298-332.jpg",
  "prompt": "A cinematic ocean wave at sunrise, highly detailed",
  "aspect_ratio": "16:9",
  "resolution": "720p",
  "duration": 8
}
JSON
)

# 1. Submit the prediction.
SUBMIT_RESPONSE=$(curl --silent --show-error --fail-with-body \
  -X POST "https://api.wavespeed.ai/api/v3/google/gemini-omni-1.1-flash/image-to-video" \
  -H "Authorization: Bearer ${WAVESPEED_API_KEY}" \
  -H "Content-Type: application/json" \
  -d "${REQUEST_BODY}")

TASK=$(printf '%s' "${SUBMIT_RESPONSE}" | jq 'if type == "object" and has("data") then .data else . end')
PREDICTION_ID=$(printf '%s' "${TASK}" | jq -r '.id // empty')
if [ -z "${PREDICTION_ID}" ]; then
  printf 'Submission response did not contain a prediction id
' >&2
  exit 1
fi
RESULT_URL="https://api.wavespeed.ai/api/v3/predictions/${PREDICTION_ID}/result"

# 2. Poll until the prediction finishes.
while true; do
  RESPONSE=$(curl --silent --show-error --fail-with-body \
    "${RESULT_URL}" \
    -H "Authorization: Bearer ${WAVESPEED_API_KEY}")
  RESULT=$(printf '%s' "${RESPONSE}" | jq 'if type == "object" and has("data") then .data else . end')
  STATUS=$(printf '%s' "${RESULT}" | jq -r '.status // empty')

  case "${STATUS}" in
    completed) printf '%s\n' "${RESULT}" | jq '.outputs'; break ;;
    failed|cancelled|timeout|deleted) printf '%s\n' "${RESULT}" | jq . >&2; exit 1 ;;
    *) sleep 2 ;;
  esac
done

Parameters

Task Submission Parameters

Request Parameters

ParameterTypeRequiredDefaultRangeDescription
imagestringYes-Input image URL to animate.
promptstringYes-Text prompt describing how the image should be animated.
last_imagestringNo--Optional end frame. When provided, the model interpolates from the first image to this image.
aspect_ratiostringNo16:916:9, 9:16Aspect ratio of the generated video.
resolutionstringNo720p360p, 720p, 1080p, 4kResolution of the generated video.
durationintegerNo83, 4, 5, 6, 7, 8, 9, 10Duration of the generated video in seconds.

Response Parameters

ParameterTypeDescription
codeintegerHTTP status code (e.g., 200 for success)
messagestringStatus message (e.g., “success”)
data.idstringUnique identifier for the prediction, Task Id
data.modelstringModel ID used for the prediction
data.outputsarrayOutput values, usually URL strings; some models return text strings or structured result objects (empty when status is not completed)
data.urlsobjectObject containing related API endpoints
data.statusstringTask status. completed is successful; failed, cancelled, timeout, and deleted are failure terminal statuses.
data.created_atstringISO timestamp of when the request was created (e.g., “2023-04-01T12:34:56.789Z”)
data.errorstringError message (empty if no error occurred)
data.timingsobjectObject containing timing details
data.timings.inferenceintegerInference time in milliseconds

Result Request Parameters

ParameterTypeRequiredDefaultDescription
idstringYes-Task ID

Result Response Parameters

ParameterTypeDescription
codeintegerHTTP status code (e.g., 200 for success)
messagestringStatus message (e.g., “success”)
dataobjectThe prediction data object containing all details
data.idstringUnique identifier for the prediction
data.modelstringModel ID used for the prediction
data.outputsarray<string | object>Array of generated outputs (empty when status is not completed). Items are usually URL strings, but may be text strings or structured result objects, depending on the model.
data.urlsobjectObject containing related API endpoints
data.statusstringStatus: completed is successful; failed, cancelled, timeout, and deleted are failure terminal statuses
data.created_atstringISO timestamp of when the request was created
data.errorstringError message (empty if no error occurred)
data.timingsobjectObject containing timing details
data.timings.inferenceintegerInference time in milliseconds
© 2026 WaveSpeedAI. All rights reserved.