GPT Image 2.5 is LIVE — Flare & Sunburst | Try in Image Generator →
Google Models

Google Models

Google's cutting-edge AI models deliver high-performance image and video models

Google's cutting-edge AI models deliver high-performance image and video models

All models

46 models
google/gemini-omni-flash/reference-to-video
image-to-video$0.1300

google/gemini-omni-flash/reference-to-video

Gemini Omni Flash Reference to Video creates short AI videos with synchronized audio from one or more reference images and a text prompt, preserving visual identity and following the provided references for guided multimodal video generation. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

google/gemini-omni-flash/video-edit
video-to-video$0.1300

google/gemini-omni-flash/video-edit

Gemini Omni Flash Video Edit applies natural-language edit instructions to existing videos, enabling prompt-guided changes to scenes, style, motion, and visual details while preserving the original video context. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

google/gemini-omni-1.1-flash/reference-to-video
image-to-video$0.1000

google/gemini-omni-1.1-flash/reference-to-video

Gemini Omni 1.1 Flash Reference-to-Video creates short AI videos with synchronized audio from a text prompt plus optional image and video references, supporting resolutions from 360P to 4K for character-consistent clips, visual reference guidance, social content, marketing videos, and production workflows. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

google/gemini-omni-1.1-flash/image-to-video
image-to-video$0.1000

google/gemini-omni-1.1-flash/image-to-video

Gemini Omni 1.1 Flash Image-to-Video animates a start image into short AI videos, optionally targeting an end frame and generating synchronized audio at resolutions from 360P to 4K for social content, creative storytelling, marketing videos, and production workflows. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

google/gemini-omni-1.1-flash/text-to-video
text-to-video$0.1000

google/gemini-omni-1.1-flash/text-to-video

Gemini Omni 1.1 Flash Text-to-Video creates short AI videos with synchronized audio from text prompts, supporting resolutions from 360P to 4K for social content, creative storytelling, marketing videos, and production workflows. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

google/nano-banana-2-lite/edit
image-to-image$0.0400

google/nano-banana-2-lite/edit

Google Nano Banana 2 Lite Edit transforms uploaded images with text instructions, supporting fast prompt-guided image editing, visual refinements, and creative changes with low latency. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

google/nano-banana-2-lite/text-to-image
text-to-image$0.0400

google/nano-banana-2-lite/text-to-image

Google Nano Banana 2 Lite Text to Image generates high-quality images from text prompts with low latency, flexible aspect ratios, and fast image creation for creative and production workflows. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

google/gemini-omni-flash/text-to-video
text-to-video$0.1300

google/gemini-omni-flash/text-to-video

Gemini Omni Flash Text to Video creates short AI videos with synchronized audio from text prompts, combining visual generation and audio output for fast multimodal video creation. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

google/gemini-omni-1.1-flash/video-edit
video-to-video$0.1200

google/gemini-omni-1.1-flash/video-edit

Gemini Omni 1.1 Flash Video Edit applies natural-language edits to existing videos, supporting output resolutions from 360P to 4K for prompt-guided video modification, scene refinements, creative edits, marketing content, and production workflows. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

google/gemini-omni-flash/image-to-video
image-to-video$0.1300

google/gemini-omni-flash/image-to-video

Gemini Omni Flash Image to Video animates input images into short AI videos with synchronized audio, adding motion and sound while following the source image. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

google/nano-banana-pro/edit10% OFF
image-to-image$0.1400$0.1260

google/nano-banana-pro/edit

Google Nano Banana Pro (Gemini 3.0 Pro Image) Edit enables image editing with 4K-capable output. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

google/nano-banana-2/edit10% OFF
image-to-image$0.0700$0.0630

google/nano-banana-2/edit

Google Nano Banana 2 Edit (Gemini 3.1 Flash Image) enables advanced image editing with 4K-capable output, fast iteration, and precise instruction following. Supports text translation, localization within images, and maintains subject consistency during edits. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

google/nano-banana-pro/edit-ultra
image-to-image$0.1500

google/nano-banana-pro/edit-ultra

Google Nano Banana Pro (Gemini 3.0 Pro Image) Edit enables image editing with highres output. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

google/nano-banana-2/edit-fast
image-to-image$0.0450

google/nano-banana-2/edit-fast

Google Nano Banana 2 Edit Fast (Gemini 3.1 Flash Image) is the cheapest Nano Banana 2 editing option, starting at just $0.045 per image. Enables fast image editing with 2K default output and 4K support. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

google/veo3.1/image-to-video
image-to-video$3.2000

google/veo3.1/image-to-video

Google Veo 3.1 is an Image-to-Video model that converts images into high-quality videos with native 1080P output for enhanced detail and creative flexibility. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

google/veo3.1-fast/image-to-video
image-to-video$1.2000

google/veo3.1-fast/image-to-video

Google Veo 3.1 Fast is an Image-to-Video model with native 1080p output for high-detail videos from images and fast performance. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

google/veo3.1-lite/image-to-video
image-to-video$0.3000

google/veo3.1-lite/image-to-video

Google Veo 3.1 Lite Image-to-Video transforms static images into high-fidelity 720p or 1080p videos with natively generated audio. Supports many interpolation use cases, landscape and portrait aspect ratios, and customizable duration. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

google/veo3.1/text-to-video
text-to-video$3.2000

google/veo3.1/text-to-video

Google Veo 3.1 converts text prompts into videos with synchronized audio at native 1080p for high-quality outputs. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

google/veo3.1-lite/start-end-to-video
image-to-video$0.4000

google/veo3.1-lite/start-end-to-video

Google Veo 3.1 Lite Start-End-to-Video generates high-fidelity videos by interpolating between a start image and an optional end image. Supports 720p and 1080p resolutions, landscape and portrait aspect ratios, and native audio generation. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

google/veo3.1/reference-to-video
image-to-video$3.2000

google/veo3.1/reference-to-video

Google Veo3.1 Reference-to-Video performs image-to-video generation that preserves a specific subject's appearance and identity from provided reference images, enabling consistent character or product motion across frames. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

google/veo3.1-fast/text-to-video
text-to-video$1.2000

google/veo3.1-fast/text-to-video

Google Veo 3.1 Fast creates text-to-video with native 1080p and synchronized audio, delivering high-quality videos for creators. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

google/veo3.1-fast/video-extend
video-extend$1.2000

google/veo3.1-fast/video-extend

Extend Veo 3.1 videos in 7-second steps with the Fast endpoint—quick, coherent continuation that preserves style and motion, output as a single merged clip. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

google/nano-banana-pro/text-to-image10% OFF
text-to-image$0.1400$0.1260

google/nano-banana-pro/text-to-image

Google's Nano Banana pro (Gemini 3.0 Pro Image) is a cutting-edge text-to-image model enabling high-res 4K image generation optimized for phones. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

google/nano-banana-pro/text-to-image-ultra
text-to-image$0.1500

google/nano-banana-pro/text-to-image-ultra

Google's Nano Banana Pro (Gemini 3.0 Pro Image) is a cutting-edge text-to-image model enabling high-res image generation optimized for phones. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

google/nano-banana-2/text-to-image10% OFF
text-to-image$0.0700$0.0630

google/nano-banana-2/text-to-image

Google Nano Banana 2 (Gemini 3.1 Flash Image) delivers Pro-quality image generation at Flash speed with 512px to 4K resolution support. Features include improved text rendering, character consistency for up to 5 characters, and real-world knowledge integration. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

google/nano-banana-pro/text-to-image-multi
text-to-image$0.0700

google/nano-banana-pro/text-to-image-multi

Google's Nano Banana Pro (Gemini 3.0 Pro Image) is a next-generation text-to-image model capable of generating multiple high-quality images in a single run. Extremely low cost — only $0.07 per image. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

google/nano-banana-2/text-to-image-fast
text-to-image$0.0450

google/nano-banana-2/text-to-image-fast

Google Nano Banana 2 Fast (Gemini 3.1 Flash Image) is the cheapest Nano Banana 2 option, starting at just $0.045 per image. Delivers fast text-to-image generation with 2K default output and 4K support. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

google/nano-banana/text-to-image
text-to-image$0.0380

google/nano-banana/text-to-image

Google Nano Banana is a cutting-edge text-to-image model that generates images from natural language prompts. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

google/nano-banana-pro/edit-multi
image-to-image$0.0700

google/nano-banana-pro/edit-multi

Google's Nano Banana Pro (Gemini 3.0 Pro Image) Edit is a next-generation image editing model capable of generating multiple high-quality edited images in a single run. Extremely low cost — only $0.07 per image. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

google/veo3.1-fast/reference-to-video
image-to-video$0.6400

google/veo3.1-fast/reference-to-video

Google Veo 3.1 Fast Reference to Video is a fast AI reference-to-video generation model that creates 8-second videos from up to three reference images using the official Veo predictLongRunning endpoint with referenceImages assets. Ready-to-use REST inference API for product videos, character consistency, branded visual storytelling, social media clips, advertising creatives, and professional reference-based video generation workflows with simple integration, no coldstarts, and affordable pricing.

google/gemini-3.1-flash/text-to-speech
text-to-audio$0.2000

google/gemini-3.1-flash/text-to-speech

Gemini 3.1 Flash Text to Speech generates expressive multi-speaker audio from text, with natural voices and multilingual language control for dialogue, narration, localization, and AI voice workflows. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

google/nano-banana/edit
image-to-image$0.0380

google/nano-banana/edit

Nano-Banana is an advanced image generation and editing model that produces photorealistic or stylized visuals and performs precise inpainting, outpainting, and background replacement. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

google/veo3
text-to-video$3.2000

google/veo3

Google Veo3 is Google's flagship text-to-video model with built-in audio, producing synchronized video and sound from text prompts. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

google/veo3-fast
text-to-video$1.2000

google/veo3-fast

Google Veo 3 Fast creates text-to-video with synchronized audio, delivering faster, more cost-effective results than standard Veo 3; commercial use allowed and pricing starts at $0.25/second. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

google/veo3-fast/image-to-video
image-to-video$1.2000

google/veo3-fast/image-to-video

Google Veo3 Fast provides faster, more cost-effective Image-to-Video generation vs Veo 3, with commercial use allowed and $0.25/sec pricing. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

google/veo3/image-to-video
image-to-video$3.2000

google/veo3/image-to-video

Google Veo 3 is Google's flagship image-to-video model that creates audio-enabled videos from images. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

google/lyria-3-pro/music
text-to-audio$0.0800

google/lyria-3-pro/music

Google Lyria 3 Pro generates high-quality music tracks from text prompts and optional image input. Pro tier delivers enhanced audio quality and richer compositions. Produces complete songs with lyrics, descriptions, and audio output. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

google/gemini-2.5-flash-image-preview/text-to-image
text-to-image$0.0380

google/gemini-2.5-flash-image-preview/text-to-image

Google Gemini 2.5 Flash Text-to-Image delivers state-of-the-art text-to-image generation and image editing with previews. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

google/gemini-2.5-flash-image/edit
image-to-image$0.0380

google/gemini-2.5-flash-image/edit

Nano Banana (Gemini 2.5 Flash Image) offers image-to-image generation and precise editing with deep reasoning for improved accuracy. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

google/gemini-2.5-flash-image/text-to-image
text-to-image$0.0380

google/gemini-2.5-flash-image/text-to-image

Google Gemini 2.5 Flash Image offers advanced text-to-image generation and image editing with creative controls for quality images. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

google/gemini-2.5-flash-image-preview/edit
image-to-image$0.0380

google/gemini-2.5-flash-image-preview/edit

Google Gemini 2.5 Flash Image Preview is an image-to-image editing model with advanced creative controls for precise image edits. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

google/veo3.1-lite/text-to-video
text-to-video$0.3000

google/veo3.1-lite/text-to-video

Google Veo 3.1 Lite Text-to-Video generates high-fidelity 720p or 1080p videos with natively generated audio from text prompts. Lightweight variant optimized for cost efficiency. Supports landscape and portrait aspect ratios, dialogue with lip-sync, and customizable duration. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

google/veo3.1/video-extend
video-extend$3.2000

google/veo3.1/video-extend

Extend and continue Veo 3.1 videos with smooth motion, preserved style, and strong scene coherence. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

google/lyria-3-clip/music
text-to-audio$0.0400

google/lyria-3-clip/music

Google Lyria 3 Clip generates novel music tracks from text prompts and optional image input. Produces complete songs with lyrics, descriptions, and audio output. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

google/gemini-2.5-flash/text-to-speech
text-to-audio$0.0200

google/gemini-2.5-flash/text-to-speech

Google Gemini 2.5 Flash Text-to-Speech delivers fast, natural multi-speaker voice synthesis with 30+ voices across 24 languages at lower cost. Perfect for dialogues, conversations, and multilingual content. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

google/gemini-2.5-pro/text-to-speech
text-to-audio$0.0400

google/gemini-2.5-pro/text-to-speech

Google Gemini 2.5 Pro Text-to-Speech delivers natural multi-speaker voice synthesis with 30+ voices across 24 languages. Perfect for dialogues, conversations, and multilingual content. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

Google Models

Google AI Models on WaveSpeedAI provide a comprehensive suite of generative AI tools for video, image, music, audio, and speech creation. The collection includes Veo for cinematic AI video generation, Imagen for high-quality image creation, Nano Banana and Nano Banana Lite for fast creative image workflows, Nano Banana Pro for premium visual production, Omni Flash for lightweight multimodal generation, Lyric for AI music creation, and Gemini Text-to-Speech for natural voice synthesis.

Built for creators, developers, marketers, and AI applications, Google AI Models support text-to-video, image-to-video, video extension, reference-based video generation, text-to-image, image editing, AI music generation, and text-to-speech workflows. These models combine strong prompt understanding, realistic motion, high visual fidelity, synchronized audio-video generation, fast iteration, and scalable API access for professional creative production.

Core Model Capabilities

Cinematic AI Video Generation:

Use Veo models to generate cinematic videos from text prompts, still images, start-end frames, or reference videos. Veo supports realistic motion, natural lighting, camera control, synchronized audio, and smooth scene continuity for storytelling, ads, product videos, and social media content.

Video Extension:

Extend existing Veo-generated videos into longer continuous clips while preserving motion style, framing, lighting, scene continuity, and synchronized audio. Fast variants are suitable for rapid previews, creative iteration, and multi-branch story continuation.

AI Image Generation:

Use Imagen, Nano Banana, Nano Banana Lite, Nano Banana Pro, Gemini, and Omni Flash models to create high-quality images from prompts for portraits, product visuals, key art, social media content, blog images, design concepts, and commercial assets.

AI Image Editing:

Transform existing images with prompt-guided editing, context-aware refinement, style adjustment, identity preservation, lighting consistency, and region-aware modifications.

Lightweight Creative Workflows:

Nano Banana, Nano Banana Lite, Gemini Flash, and Omni Flash provide faster and more cost-efficient generation options for everyday creative tasks, rapid prototyping, preview workflows, and high-volume image production.

Premium Image Workflows:

Imagen 4 Ultra, Nano Banana Pro, Nano Banana Pro Ultra, and Nano Banana Pro Multi support higher-fidelity image generation, improved prompt control, multi-reference consistency, and premium visual output for hero shots, advertising campaigns, brand assets, and professional creative projects.

AI Music Generation:

Lyric models generate high-quality music from prompts, supporting background scoring, social media content, soundtrack creation, and professional audio production workflows.

Text-to-Speech:

Gemini Text-to-Speech models provide natural and expressive voice synthesis for narration, dialogue, avatars, education, product explainers, and multilingual audio content.

Google AI Models on WaveSpeedAI give creators and developers fast access to Google's video, image, music, audio, and speech generation models with scalable APIs, flexible pricing, and production-ready creative capabilities.

Google Models API — pricing & performance

Run any model in the Google Models collection through a single REST API. Pay per generation — no subscriptions, no minimums — with industry-leading latency on a 99.9% uptime infrastructure.

Why run Google Models on WaveSpeedAI

Transparent pricing

Per-call pricing for every Google Models model. The price is listed on each model page — no platform fees on top.

Optimized for low latency

Most Google Models image models complete in under 2 seconds. Video and 3D models run several times faster than self-hosted alternatives.

99.9% uptime

Multi-region failover and automatic retries keep your production traffic online — even during provider outages.

Frequently asked questions

How much does the Google Models API cost?+

Each model has its own per-call price listed on the model page. We bill per successful generation, with no subscription fees or minimums.

How fast are Google Models models on WaveSpeedAI?+

Image models in this collection typically complete in under 2 seconds. Video and 3D models depend on duration and resolution but are usually several times faster than self-hosted runs.

Can I try the API without a credit card?+

Eligible new accounts may receive $1 in promotional credits to try Google Models models without a credit card. Trial credits are not guaranteed for every signup; check your account balance before generating.

Are there rate limits?+

Standard accounts have generous concurrent-job limits. Enterprise plans offer custom RPM, higher concurrency, and dedicated capacity — contact sales for details.