Google Gemini Omni 1.1 Flash Text To Video API Documentation

Google Gemini Omni 1.1 Flash Text To Video API Documentation

Playground

Try it on WaveSpeedAI!

Gemini Omni 1.1 Flash Text-to-Video creates short AI videos with synchronized audio from text prompts, supporting resolutions from 360P to 4K for social content, creative storytelling, marketing videos, and production workflows. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

Features

Gemini Omni 1.1 Flash Text-to-Video turns text prompts into videos with synchronized audio. It supports fast 360p drafts for prompt exploration, standard 720p generation, and higher-resolution 1080p or 4k output for delivery-ready results.


Why Choose This?

  • Text-to-video generation
    Generate videos directly from natural-language prompts.

  • Synchronized audio
    Create motion and matching audio together from one prompt.

  • Fast draft workflow
    Use 360p to test scene ideas, camera movement, timing, and prompt variations quickly.

  • Flexible resolution options
    Choose 360p, 720p, 1080p, or 4k depending on speed, cost, and output quality needs.

  • Landscape and vertical formats
    Use 16:9 for widescreen video or 9:16 for vertical content.

  • Directable creative control
    Describe framing, camera movement, scene progression, dialogue, music, ambience, and visual style in the prompt.


Parameters

ParameterRequiredDescription
promptYesScene, motion, camera movement, style, dialogue, music, ambience, and audio instructions.
aspect_ratioNoOutput aspect ratio: 16:9 or 9:16. Default: 16:9.
resolutionNoOutput resolution: 360p, 720p, 1080p, or 4k. Default: 720p.
durationNoVideo duration in seconds. Range: 3–10. Default: 8.

How to Use

  1. Write your prompt — Describe the subject, action, scene, camera movement, visual style, and audio direction.
  2. Choose aspect ratio — Use 16:9 for landscape video or 9:16 for vertical content.
  3. Select resolution — Use 360p for drafts, 720p for standard generation, or 1080p / 4k for higher-quality output.
  4. Set duration — Choose a duration from 3 to 10 seconds.
  5. Submit — Generate the final video with synchronized audio.

Pricing

Pricing is based on generated video duration and selected resolution.

ResolutionPrice per second
360p$0.03
720p$0.10
1080p$0.15
4k$0.30

Example Costs

Resolution5s8s10s
360p$0.15$0.24$0.30
720p$0.50$0.80$1.00
1080p$0.75$1.20$1.50
4k$1.50$2.40$3.00

Best Use Cases

  • Prompt exploration — Use 360p to test creative direction before generating higher-resolution versions.
  • Cinematic short clips — Create short scenes with motion, camera direction, and synchronized sound.
  • Social media videos — Generate vertical or landscape videos for posts, ads, shorts, and reels.
  • Marketing concepts — Produce campaign clips, product scenes, and promotional video ideas.
  • Storyboarding and previsualization — Turn scene descriptions into audiovisual previews.
  • High-resolution delivery — Use 1080p or 4k when final output quality matters.

Pro Tips

  • Start with 360p for fast prompt testing, then move to 1080p or 4k for final output.
  • Include subject, action, camera movement, lighting, mood, and audio cues in the prompt.
  • Use 9:16 for mobile-first video and 16:9 for widescreen content.
  • Keep the scene focused when generating short clips.
  • Mention dialogue, music, ambience, or sound effects directly when audio direction matters.
  • Use shorter durations for quick iteration and longer durations when the scene needs more time to develop.

Authentication

For authentication details, please refer to the Authentication Guide.

API Endpoints

Submit Task & Query Result

set -euo pipefail

export WAVESPEED_API_KEY="your-api-key"

REQUEST_BODY=$(cat <<'JSON'
{
  "prompt": "A cinematic ocean wave at sunrise, highly detailed",
  "aspect_ratio": "16:9",
  "resolution": "720p",
  "duration": 8
}
JSON
)

# 1. Submit the prediction.
SUBMIT_RESPONSE=$(curl --silent --show-error --fail-with-body \
  -X POST "https://api.wavespeed.ai/api/v3/google/gemini-omni-1.1-flash/text-to-video" \
  -H "Authorization: Bearer ${WAVESPEED_API_KEY}" \
  -H "Content-Type: application/json" \
  -d "${REQUEST_BODY}")

TASK=$(printf '%s' "${SUBMIT_RESPONSE}" | jq 'if type == "object" and has("data") then .data else . end')
PREDICTION_ID=$(printf '%s' "${TASK}" | jq -r '.id // empty')
if [ -z "${PREDICTION_ID}" ]; then
  printf 'Submission response did not contain a prediction id
' >&2
  exit 1
fi
RESULT_URL="https://api.wavespeed.ai/api/v3/predictions/${PREDICTION_ID}/result"

# 2. Poll until the prediction finishes.
while true; do
  RESPONSE=$(curl --silent --show-error --fail-with-body \
    "${RESULT_URL}" \
    -H "Authorization: Bearer ${WAVESPEED_API_KEY}")
  RESULT=$(printf '%s' "${RESPONSE}" | jq 'if type == "object" and has("data") then .data else . end')
  STATUS=$(printf '%s' "${RESULT}" | jq -r '.status // empty')

  case "${STATUS}" in
    completed) printf '%s\n' "${RESULT}" | jq '.outputs'; break ;;
    failed|cancelled|timeout|deleted) printf '%s\n' "${RESULT}" | jq . >&2; exit 1 ;;
    *) sleep 2 ;;
  esac
done

Parameters

Task Submission Parameters

Request Parameters

ParameterTypeRequiredDefaultRangeDescription
promptstringYes-Text prompt describing the video to generate.
aspect_ratiostringNo16:916:9, 9:16Aspect ratio of the generated video.
resolutionstringNo720p360p, 720p, 1080p, 4kResolution of the generated video.
durationintegerNo83, 4, 5, 6, 7, 8, 9, 10Duration of the generated video in seconds.

Response Parameters

ParameterTypeDescription
codeintegerHTTP status code (e.g., 200 for success)
messagestringStatus message (e.g., “success”)
data.idstringUnique identifier for the prediction, Task Id
data.modelstringModel ID used for the prediction
data.outputsarrayOutput values, usually URL strings; some models return text strings or structured result objects (empty when status is not completed)
data.urlsobjectObject containing related API endpoints
data.statusstringTask status. completed is successful; failed, cancelled, timeout, and deleted are failure terminal statuses.
data.created_atstringISO timestamp of when the request was created (e.g., “2023-04-01T12:34:56.789Z”)
data.errorstringError message (empty if no error occurred)
data.timingsobjectObject containing timing details
data.timings.inferenceintegerInference time in milliseconds

Result Request Parameters

ParameterTypeRequiredDefaultDescription
idstringYes-Task ID

Result Response Parameters

ParameterTypeDescription
codeintegerHTTP status code (e.g., 200 for success)
messagestringStatus message (e.g., “success”)
dataobjectThe prediction data object containing all details
data.idstringUnique identifier for the prediction
data.modelstringModel ID used for the prediction
data.outputsarray<string | object>Array of generated outputs (empty when status is not completed). Items are usually URL strings, but may be text strings or structured result objects, depending on the model.
data.urlsobjectObject containing related API endpoints
data.statusstringStatus: completed is successful; failed, cancelled, timeout, and deleted are failure terminal statuses
data.created_atstringISO timestamp of when the request was created
data.errorstringError message (empty if no error occurred)
data.timingsobjectObject containing timing details
data.timings.inferenceintegerInference time in milliseconds
© 2026 WaveSpeedAI. All rights reserved.