Google Gemini Omni 1.1 Flash Reference To Video API Documentation

Google Gemini Omni 1.1 Flash Reference To Video API Documentation

Playground

Try it on WaveSpeedAI!

Gemini Omni 1.1 Flash Reference-to-Video creates short AI videos with synchronized audio from a text prompt plus optional image and video references, supporting resolutions from 360P to 4K for character-consistent clips, visual reference guidance, social content, marketing videos, and production workflows. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

Features

Gemini Omni 1.1 Flash Reference-to-Video generates synchronized audio and video from a prompt plus reference media. Use reference images for subject, character, product, or environment guidance, and add short reference videos when motion, camera behavior, or scene rhythm matters.


Why Choose This?

  • Multimodal reference guidance
    Combine reference images and short reference videos to guide the final result.

  • Character and context consistency
    Preserve recognizable subjects, visual identity, and scene context while generating a new video.

  • Motion and camera guidance
    Use reference videos to communicate gestures, action, pacing, or camera movement that may be difficult to describe with text alone.

  • Synchronized audio
    Generate video and matching audio together in one workflow.

  • Draft-to-production workflow
    Start with 360p for faster testing, then move to 1080p or 4k for higher-quality delivery.

  • Flexible output formats
    Supports both 16:9 and 9:16 for landscape and vertical video creation.


Parameters

ParameterRequiredDescription
promptYesDescribe the desired scene and explain how the references should be used.
imagesNoUp to 10 reference image URLs.
reference_videosNoUp to 3 reference video URLs. Each reference video must be no longer than 3 seconds.
aspect_ratioNoOutput aspect ratio: 16:9 or 9:16. Default: 16:9.
resolutionNoOutput resolution: 360p, 720p, 1080p, or 4k. Default: 720p.
durationNoVideo duration in seconds. Range: 3–10. Default: 8.

How to Use

  1. Add reference media — Upload reference images, reference videos, or both.
  2. Write your prompt — Describe the target scene, action, camera behavior, and the role of each reference.
  3. Choose aspect ratio — Use 16:9 for landscape video or 9:16 for vertical content.
  4. Select resolution — Use 360p for quick testing, 720p for standard output, or 1080p / 4k for higher-quality delivery.
  5. Set duration — Choose a duration from 3 to 10 seconds.
  6. Submit — Generate the final video with synchronized audio.

Pricing

Pricing is based on generated video duration and selected resolution.

ResolutionPrice per second
360p$0.03
720p$0.10
1080p$0.15
4k$0.30

Example Costs

Resolution5s8s10s
360p$0.15$0.24$0.30
720p$0.50$0.80$1.00
1080p$0.75$1.20$1.50
4k$1.50$2.40$3.00

Best Use Cases

  • Reference-guided video generation — Create videos that follow specific people, products, objects, or environments.
  • Character-consistent scenes — Keep visual identity more stable across generated shots.
  • Motion transfer and camera guidance — Use reference videos for action, gesture, pacing, or shot behavior.
  • Marketing and social content — Combine brand visuals, products, and motion references into short-form videos.
  • Creative prototyping — Quickly test prompt-plus-reference workflows before final production.

Pro Tips

  • Use clear, relevant reference images when identity or appearance matters.
  • Use reference videos when motion, pacing, or camera behavior is important.
  • In the prompt, clearly describe what each reference should influence.
  • Avoid mixing too many conflicting references in one request.
  • Start at 360p for iteration, then re-run at a higher resolution for final output.
  • Keep reference videos short and focused on the motion you want to guide.

Notes

  • prompt is required.
  • At least one reference input should be provided through images or reference_videos.
  • images supports up to 10 inputs.
  • reference_videos supports up to 3 inputs, and each video must be at most 3 seconds.
  • The model generates synchronized audio together with the video.

Authentication

For authentication details, please refer to the Authentication Guide.

API Endpoints

Submit Task & Query Result

set -euo pipefail

export WAVESPEED_API_KEY="your-api-key"

REQUEST_BODY=$(cat <<'JSON'
{
  "prompt": "A cinematic ocean wave at sunrise, highly detailed",
  "aspect_ratio": "16:9",
  "resolution": "720p",
  "duration": 8
}
JSON
)

# 1. Submit the prediction.
SUBMIT_RESPONSE=$(curl --silent --show-error --fail-with-body \
  -X POST "https://api.wavespeed.ai/api/v3/google/gemini-omni-1.1-flash/reference-to-video" \
  -H "Authorization: Bearer ${WAVESPEED_API_KEY}" \
  -H "Content-Type: application/json" \
  -d "${REQUEST_BODY}")

TASK=$(printf '%s' "${SUBMIT_RESPONSE}" | jq 'if type == "object" and has("data") then .data else . end')
PREDICTION_ID=$(printf '%s' "${TASK}" | jq -r '.id // empty')
if [ -z "${PREDICTION_ID}" ]; then
  printf 'Submission response did not contain a prediction id
' >&2
  exit 1
fi
RESULT_URL="https://api.wavespeed.ai/api/v3/predictions/${PREDICTION_ID}/result"

# 2. Poll until the prediction finishes.
while true; do
  RESPONSE=$(curl --silent --show-error --fail-with-body \
    "${RESULT_URL}" \
    -H "Authorization: Bearer ${WAVESPEED_API_KEY}")
  RESULT=$(printf '%s' "${RESPONSE}" | jq 'if type == "object" and has("data") then .data else . end')
  STATUS=$(printf '%s' "${RESULT}" | jq -r '.status // empty')

  case "${STATUS}" in
    completed) printf '%s\n' "${RESULT}" | jq '.outputs'; break ;;
    failed|cancelled|timeout|deleted) printf '%s\n' "${RESULT}" | jq . >&2; exit 1 ;;
    *) sleep 2 ;;
  esac
done

Parameters

Task Submission Parameters

Request Parameters

ParameterTypeRequiredDefaultRangeDescription
promptstringYes-Text prompt describing the video and how the reference media should guide the result.
imagesarray<string>No-0 ~ 10 itemsReference image URLs to incorporate into the video.
reference_videosarray<string>No-0 ~ 3 itemsUp to three reference video URLs. Each video must be no longer than 3 seconds.
aspect_ratiostringNo16:916:9, 9:16Aspect ratio of the generated video.
resolutionstringNo720p360p, 720p, 1080p, 4kResolution of the generated video.
durationintegerNo83, 4, 5, 6, 7, 8, 9, 10Duration of the generated video in seconds.

Response Parameters

ParameterTypeDescription
codeintegerHTTP status code (e.g., 200 for success)
messagestringStatus message (e.g., “success”)
data.idstringUnique identifier for the prediction, Task Id
data.modelstringModel ID used for the prediction
data.outputsarrayOutput values, usually URL strings; some models return text strings or structured result objects (empty when status is not completed)
data.urlsobjectObject containing related API endpoints
data.statusstringTask status. completed is successful; failed, cancelled, timeout, and deleted are failure terminal statuses.
data.created_atstringISO timestamp of when the request was created (e.g., “2023-04-01T12:34:56.789Z”)
data.errorstringError message (empty if no error occurred)
data.timingsobjectObject containing timing details
data.timings.inferenceintegerInference time in milliseconds

Result Request Parameters

ParameterTypeRequiredDefaultDescription
idstringYes-Task ID

Result Response Parameters

ParameterTypeDescription
codeintegerHTTP status code (e.g., 200 for success)
messagestringStatus message (e.g., “success”)
dataobjectThe prediction data object containing all details
data.idstringUnique identifier for the prediction
data.modelstringModel ID used for the prediction
data.outputsarray<string | object>Array of generated outputs (empty when status is not completed). Items are usually URL strings, but may be text strings or structured result objects, depending on the model.
data.urlsobjectObject containing related API endpoints
data.statusstringStatus: completed is successful; failed, cancelled, timeout, and deleted are failure terminal statuses
data.created_atstringISO timestamp of when the request was created
data.errorstringError message (empty if no error occurred)
data.timingsobjectObject containing timing details
data.timings.inferenceintegerInference time in milliseconds
© 2026 WaveSpeedAI. All rights reserved.