Google Gemini Omni 1.1 Flash Image To Video API Documentation
Playground
Try it on WaveSpeedAI!Gemini Omni 1.1 Flash Image-to-Video animates a start image into short AI videos, optionally targeting an end frame and generating synchronized audio at resolutions from 360P to 4K for social content, creative storytelling, marketing videos, and production workflows. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
Features
Gemini Omni 1.1 Flash Image-to-Video animates a start image into a video with synchronized audio. You can optionally add a final frame to guide the ending of the shot, making it useful for controlled transitions, reveal sequences, camera movement, and other prompt-guided image-to-video workflows.
Why Choose This?
-
Start-frame image animation
Turn a single image into a video with motion, scene progression, and synchronized audio. -
Optional end-frame guidance
Addlast_imagewhen you want more control over how the shot ends. -
Smooth transition generation
Useful for camera moves, reveal shots, visual transformations, and controlled motion between two keyframes. -
Draft-to-production workflow
Use360pfor quick testing, then move to1080por4kfor higher-quality output. -
Synchronized audio
Generate motion and matching audio together in one workflow. -
Flexible output formats
Supports both16:9and9:16for landscape and vertical video creation.
Parameters
| Parameter | Required | Description |
|---|---|---|
| image | Yes | Start-frame image URL. |
| last_image | No | Optional end-frame image URL used to guide the ending of the video. |
| prompt | Yes | Motion, camera movement, transition, scene, style, and audio instructions. |
| aspect_ratio | No | Output aspect ratio: 16:9 or 9:16. Default: 16:9. |
| resolution | No | Output resolution: 360p, 720p, 1080p, or 4k. Default: 720p. |
| duration | No | Video duration in seconds. Range: 3–10. Default: 8. |
How to Use
- Upload the start image — Provide the image that should define the first frame.
- Add an end image optional — Use
last_imagewhen you want to guide the final frame or overall transition direction. - Write your prompt — Describe the motion, camera movement, transition behavior, scene changes, style, and audio.
- Choose aspect ratio — Use
16:9for landscape video or9:16for vertical content. - Select resolution — Use
360pfor drafts,720pfor standard output, or1080p/4kfor higher-quality results. - Set duration — Choose a duration from
3to10seconds. - Submit — Generate the final video with synchronized audio.
Pricing
Pricing is based on generated video duration and selected resolution.
| Resolution | Price per second |
|---|---|
| 360p | $0.03 |
| 720p | $0.10 |
| 1080p | $0.15 |
| 4k | $0.30 |
Example Costs
| Resolution | 5s | 8s | 10s |
|---|---|---|---|
| 360p | $0.15 | $0.24 | $0.30 |
| 720p | $0.50 | $0.80 | $1.00 |
| 1080p | $0.75 | $1.20 | $1.50 |
| 4k | $1.50 | $2.40 | $3.00 |
Best Use Cases
- Image animation — Turn still images into short motion clips with audio.
- Controlled transitions — Use
imageandlast_imageto guide the beginning and end of a shot. - Reveal and transformation scenes — Create visual change between two defined states.
- Social media content — Generate vertical or landscape motion content from static artwork or product images.
- Creative prototyping — Quickly test motion, pacing, and shot progression from a single keyframe.
- Marketing and campaign assets — Animate posters, product visuals, and branded images into short videos.
Pro Tips
- Use a clear, high-quality start image for better subject consistency.
- Add
last_imagewhen the final pose, composition, or scene state matters. - Describe both what moves and how it moves in the prompt.
- Include camera instructions such as push-in, orbit, pan, or pull-back for more controlled results.
- Mention ambience, music, dialogue, or sound effects directly when audio guidance matters.
- Start with
360pfor fast iteration, then move to1080por4kfor final output. - Use shorter durations for quick tests and longer durations when the transition needs more time to develop.
Related Models
- Gemini Omni 1.1 Flash Text-to-Video — Generate videos directly from text prompts with synchronized audio.
- Gemini Omni 1.1 Flash Reference-to-Video — Generate videos using reference media for stronger visual guidance.
- Gemini Omni 1.1 Flash Video Edit — Edit existing videos with prompt instructions.
Authentication
For authentication details, please refer to the Authentication Guide.
API Endpoints
Submit Task & Query Result
set -euo pipefail
export WAVESPEED_API_KEY="your-api-key"
REQUEST_BODY=$(cat <<'JSON'
{
"image": "https://interactive-examples.mdn.mozilla.net/media/cc0-images/painted-hand-298-332.jpg",
"prompt": "A cinematic ocean wave at sunrise, highly detailed",
"aspect_ratio": "16:9",
"resolution": "720p",
"duration": 8
}
JSON
)
# 1. Submit the prediction.
SUBMIT_RESPONSE=$(curl --silent --show-error --fail-with-body \
-X POST "https://api.wavespeed.ai/api/v3/google/gemini-omni-1.1-flash/image-to-video" \
-H "Authorization: Bearer ${WAVESPEED_API_KEY}" \
-H "Content-Type: application/json" \
-d "${REQUEST_BODY}")
TASK=$(printf '%s' "${SUBMIT_RESPONSE}" | jq 'if type == "object" and has("data") then .data else . end')
PREDICTION_ID=$(printf '%s' "${TASK}" | jq -r '.id // empty')
if [ -z "${PREDICTION_ID}" ]; then
printf 'Submission response did not contain a prediction id
' >&2
exit 1
fi
RESULT_URL="https://api.wavespeed.ai/api/v3/predictions/${PREDICTION_ID}/result"
# 2. Poll until the prediction finishes.
while true; do
RESPONSE=$(curl --silent --show-error --fail-with-body \
"${RESULT_URL}" \
-H "Authorization: Bearer ${WAVESPEED_API_KEY}")
RESULT=$(printf '%s' "${RESPONSE}" | jq 'if type == "object" and has("data") then .data else . end')
STATUS=$(printf '%s' "${RESULT}" | jq -r '.status // empty')
case "${STATUS}" in
completed) printf '%s\n' "${RESULT}" | jq '.outputs'; break ;;
failed|cancelled|timeout|deleted) printf '%s\n' "${RESULT}" | jq . >&2; exit 1 ;;
*) sleep 2 ;;
esac
doneParameters
Task Submission Parameters
Request Parameters
| Parameter | Type | Required | Default | Range | Description |
|---|---|---|---|---|---|
| image | string | Yes | - | Input image URL to animate. | |
| prompt | string | Yes | - | Text prompt describing how the image should be animated. | |
| last_image | string | No | - | - | Optional end frame. When provided, the model interpolates from the first image to this image. |
| aspect_ratio | string | No | 16:9 | 16:9, 9:16 | Aspect ratio of the generated video. |
| resolution | string | No | 720p | 360p, 720p, 1080p, 4k | Resolution of the generated video. |
| duration | integer | No | 8 | 3, 4, 5, 6, 7, 8, 9, 10 | Duration of the generated video in seconds. |
Response Parameters
| Parameter | Type | Description |
|---|---|---|
| code | integer | HTTP status code (e.g., 200 for success) |
| message | string | Status message (e.g., “success”) |
| data.id | string | Unique identifier for the prediction, Task Id |
| data.model | string | Model ID used for the prediction |
| data.outputs | array | Output values, usually URL strings; some models return text strings or structured result objects (empty when status is not completed) |
| data.urls | object | Object containing related API endpoints |
| data.status | string | Task status. completed is successful; failed, cancelled, timeout, and deleted are failure terminal statuses. |
| data.created_at | string | ISO timestamp of when the request was created (e.g., “2023-04-01T12:34:56.789Z”) |
| data.error | string | Error message (empty if no error occurred) |
| data.timings | object | Object containing timing details |
| data.timings.inference | integer | Inference time in milliseconds |
Result Request Parameters
| Parameter | Type | Required | Default | Description |
|---|---|---|---|---|
| id | string | Yes | - | Task ID |
Result Response Parameters
| Parameter | Type | Description |
|---|---|---|
| code | integer | HTTP status code (e.g., 200 for success) |
| message | string | Status message (e.g., “success”) |
| data | object | The prediction data object containing all details |
| data.id | string | Unique identifier for the prediction |
| data.model | string | Model ID used for the prediction |
| data.outputs | array<string | object> | Array of generated outputs (empty when status is not completed). Items are usually URL strings, but may be text strings or structured result objects, depending on the model. |
| data.urls | object | Object containing related API endpoints |
| data.status | string | Status: completed is successful; failed, cancelled, timeout, and deleted are failure terminal statuses |
| data.created_at | string | ISO timestamp of when the request was created |
| data.error | string | Error message (empty if no error occurred) |
| data.timings | object | Object containing timing details |
| data.timings.inference | integer | Inference time in milliseconds |