MiniMax H3 开源权重 | 在视频生成器中体验 →
Google Models

Google Models

Google's cutting-edge AI models deliver high-performance image and video models

Google's cutting-edge AI models deliver high-performance image and video models

所有模型

47 个模型
google/gemini-omni-flash/text-to-video
text-to-video

google/gemini-omni-flash/text-to-video

Gemini Omni Flash Text to Video creates short AI videos with synchronized audio from text prompts, combining visual generation and audio output for fast multimodal video creation. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

google/nano-banana-2-lite/edit
image-to-image

google/nano-banana-2-lite/edit

Google Nano Banana 2 Lite Edit transforms uploaded images with text instructions, supporting fast prompt-guided image editing, visual refinements, and creative changes with low latency. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

google/nano-banana-2-lite/text-to-image
text-to-image

google/nano-banana-2-lite/text-to-image

Google Nano Banana 2 Lite Text to Image generates high-quality images from text prompts with low latency, flexible aspect ratios, and fast image creation for creative and production workflows. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

google/gemini-omni-flash/video-edit
video-to-video

google/gemini-omni-flash/video-edit

Gemini Omni Flash Video Edit applies natural-language edit instructions to existing videos, enabling prompt-guided changes to scenes, style, motion, and visual details while preserving the original video context. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

google/gemini-omni-flash/reference-to-video
image-to-video

google/gemini-omni-flash/reference-to-video

Gemini Omni Flash Reference to Video creates short AI videos with synchronized audio from one or more reference images and a text prompt, preserving visual identity and following the provided references for guided multimodal video generation. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

google/gemini-omni-flash/image-to-video
image-to-video

google/gemini-omni-flash/image-to-video

Gemini Omni Flash Image to Video animates input images into short AI videos with synchronized audio, adding motion and sound while following the source image. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

google/nano-banana-pro/edit
image-to-image

google/nano-banana-pro/edit

Google Nano Banana Pro (Gemini 3.0 Pro Image) Edit enables image editing with 4K-capable output. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

google/nano-banana-2/edit
image-to-image

google/nano-banana-2/edit

Google Nano Banana 2 Edit (Gemini 3.1 Flash Image) enables advanced image editing with 4K-capable output, fast iteration, and precise instruction following. Supports text translation, localization within images, and maintains subject consistency during edits. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

google/nano-banana-pro/edit-ultra
image-to-image

google/nano-banana-pro/edit-ultra

Google Nano Banana Pro (Gemini 3.0 Pro Image) Edit enables image editing with highres output. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

google/nano-banana-2/edit-fast
image-to-image

google/nano-banana-2/edit-fast

Google Nano Banana 2 Edit Fast (Gemini 3.1 Flash Image) is the cheapest Nano Banana 2 editing option, starting at just $0.045 per image. Enables fast image editing with 2K default output and 4K support. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

google/veo3.1/image-to-video
image-to-video

google/veo3.1/image-to-video

Google Veo 3.1 is an Image-to-Video model that converts images into high-quality videos with native 1080P output for enhanced detail and creative flexibility. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

google/veo3.1-fast/image-to-video
image-to-video

google/veo3.1-fast/image-to-video

Google Veo 3.1 Fast is an Image-to-Video model with native 1080p output for high-detail videos from images and fast performance. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

google/veo3.1-lite/image-to-video
image-to-video

google/veo3.1-lite/image-to-video

Google Veo 3.1 Lite Image-to-Video transforms static images into high-fidelity 720p or 1080p videos with natively generated audio. Supports many interpolation use cases, landscape and portrait aspect ratios, and customizable duration. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

google/veo3.1/text-to-video
text-to-video

google/veo3.1/text-to-video

Google Veo 3.1 converts text prompts into videos with synchronized audio at native 1080p for high-quality outputs. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

google/veo3.1-lite/start-end-to-video
image-to-video

google/veo3.1-lite/start-end-to-video

Google Veo 3.1 Lite Start-End-to-Video generates high-fidelity videos by interpolating between a start image and an optional end image. Supports 720p and 1080p resolutions, landscape and portrait aspect ratios, and native audio generation. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

google/veo3.1/reference-to-video
image-to-video

google/veo3.1/reference-to-video

Google Veo3.1 Reference-to-Video performs image-to-video generation that preserves a specific subject's appearance and identity from provided reference images, enabling consistent character or product motion across frames. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

google/veo3.1-fast/text-to-video
text-to-video

google/veo3.1-fast/text-to-video

Google Veo 3.1 Fast creates text-to-video with native 1080p and synchronized audio, delivering high-quality videos for creators. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

google/veo3.1-fast/video-extend
video-extend

google/veo3.1-fast/video-extend

Extend Veo 3.1 videos in 7-second steps with the Fast endpoint—quick, coherent continuation that preserves style and motion, output as a single merged clip. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

google/nano-banana-pro/text-to-image
text-to-image

google/nano-banana-pro/text-to-image

Google's Nano Banana pro (Gemini 3.0 Pro Image) is a cutting-edge text-to-image model enabling high-res 4K image generation optimized for phones. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

google/nano-banana-pro/text-to-image-ultra
text-to-image

google/nano-banana-pro/text-to-image-ultra

Google's Nano Banana Pro (Gemini 3.0 Pro Image) is a cutting-edge text-to-image model enabling high-res image generation optimized for phones. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

google/nano-banana/edit
image-to-image

google/nano-banana/edit

Nano-Banana is an advanced image generation and editing model that produces photorealistic or stylized visuals and performs precise inpainting, outpainting, and background replacement. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

google/nano-banana-2/text-to-image
text-to-image

google/nano-banana-2/text-to-image

Google Nano Banana 2 (Gemini 3.1 Flash Image) delivers Pro-quality image generation at Flash speed with 512px to 4K resolution support. Features include improved text rendering, character consistency for up to 5 characters, and real-world knowledge integration. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

google/nano-banana-pro/text-to-image-multi
text-to-image

google/nano-banana-pro/text-to-image-multi

Google's Nano Banana Pro (Gemini 3.0 Pro Image) is a next-generation text-to-image model capable of generating multiple high-quality images in a single run. Extremely low cost — only $0.07 per image. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

google/nano-banana-2/text-to-image-fast
text-to-image

google/nano-banana-2/text-to-image-fast

Google Nano Banana 2 Fast (Gemini 3.1 Flash Image) is the cheapest Nano Banana 2 option, starting at just $0.045 per image. Delivers fast text-to-image generation with 2K default output and 4K support. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

google/nano-banana/text-to-image
text-to-image

google/nano-banana/text-to-image

Google Nano Banana is a cutting-edge text-to-image model that generates images from natural language prompts. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

google/nano-banana-pro/edit-multi
image-to-image

google/nano-banana-pro/edit-multi

Google's Nano Banana Pro (Gemini 3.0 Pro Image) Edit is a next-generation image editing model capable of generating multiple high-quality edited images in a single run. Extremely low cost — only $0.07 per image. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

google/veo3.1-fast/reference-to-video
image-to-video

google/veo3.1-fast/reference-to-video

Google Veo 3.1 Fast Reference to Video is a fast AI reference-to-video generation model that creates 8-second videos from up to three reference images using the official Veo predictLongRunning endpoint with referenceImages assets. Ready-to-use REST inference API for product videos, character consistency, branded visual storytelling, social media clips, advertising creatives, and professional reference-based video generation workflows with simple integration, no coldstarts, and affordable pricing.

google/gemini-3.1-flash/text-to-speech
text-to-audio

google/gemini-3.1-flash/text-to-speech

Gemini 3.1 Flash Text to Speech generates expressive multi-speaker audio from text, with natural voices and multilingual language control for dialogue, narration, localization, and AI voice workflows. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

google/imagen4
text-to-image

google/imagen4

Google's Imagen 4 is the flagship text-to-image model for generating images from text prompts with strong fidelity and creative control. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

google/imagen4-ultra
text-to-image

google/imagen4-ultra

Imagen4 Ultra is Google's highest-quality text-to-image model, generating high-fidelity images from simple text prompts. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

google/imagen4-fast
text-to-image

google/imagen4-fast

Google Imagen4 Fast is the fast variant of Google's Imagen 4 flagship text-to-image model for high-quality image generation. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

google/imagen3
text-to-image

google/imagen3

Imagen3 is Google's highest-quality text-to-image model, generating highly detailed, beautifully lit and photoreal images from text prompts. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

google/imagen3-fast
text-to-image

google/imagen3-fast

Imagen3 Fast is Google's top text-to-image model, creating richly detailed, beautifully lit images. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

google/veo3
text-to-video

google/veo3

Google Veo3 is Google's flagship text-to-video model with built-in audio, producing synchronized video and sound from text prompts. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

google/veo3-fast
text-to-video

google/veo3-fast

Google Veo 3 Fast creates text-to-video with synchronized audio, delivering faster, more cost-effective results than standard Veo 3; commercial use allowed and pricing starts at $0.25/second. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

google/veo3-fast/image-to-video
image-to-video

google/veo3-fast/image-to-video

Google Veo3 Fast provides faster, more cost-effective Image-to-Video generation vs Veo 3, with commercial use allowed and $0.25/sec pricing. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

google/veo3/image-to-video
image-to-video

google/veo3/image-to-video

Google Veo 3 is Google's flagship image-to-video model that creates audio-enabled videos from images. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

google/gemini-2.5-flash/text-to-speech
text-to-audio

google/gemini-2.5-flash/text-to-speech

Google Gemini 2.5 Flash Text-to-Speech delivers fast, natural multi-speaker voice synthesis with 30+ voices across 24 languages at lower cost. Perfect for dialogues, conversations, and multilingual content. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

google/veo3.1-lite/text-to-video
text-to-video

google/veo3.1-lite/text-to-video

Google Veo 3.1 Lite Text-to-Video generates high-fidelity 720p or 1080p videos with natively generated audio from text prompts. Lightweight variant optimized for cost efficiency. Supports landscape and portrait aspect ratios, dialogue with lip-sync, and customizable duration. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

google/lyria-3-clip/music
text-to-audio

google/lyria-3-clip/music

Google Lyria 3 Clip generates novel music tracks from text prompts and optional image input. Produces complete songs with lyrics, descriptions, and audio output. Supports negative prompts and seed control for reproducible results. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

google/lyria-3-pro/music
text-to-audio

google/lyria-3-pro/music

Google Lyria 3 Pro generates high-quality music tracks from text prompts and optional image input. Pro tier delivers enhanced audio quality and richer compositions. Produces complete songs with lyrics, descriptions, and audio output. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

google/gemini-2.5-flash-image/text-to-image
text-to-image

google/gemini-2.5-flash-image/text-to-image

Google Gemini 2.5 Flash Image offers advanced text-to-image generation and image editing with creative controls for quality images. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

google/gemini-2.5-flash-image/edit
image-to-image

google/gemini-2.5-flash-image/edit

Nano Banana (Gemini 2.5 Flash Image) offers image-to-image generation and precise editing with deep reasoning for improved accuracy. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

google/gemini-2.5-flash-image-preview/edit
image-to-image

google/gemini-2.5-flash-image-preview/edit

Google Gemini 2.5 Flash Image Preview is an image-to-image editing model with advanced creative controls for precise image edits. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

google/gemini-2.5-pro/text-to-speech
text-to-audio

google/gemini-2.5-pro/text-to-speech

Google Gemini 2.5 Pro Text-to-Speech delivers natural multi-speaker voice synthesis with 30+ voices across 24 languages. Perfect for dialogues, conversations, and multilingual content. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

google/veo3.1/video-extend
video-extend

google/veo3.1/video-extend

Extend and continue Veo 3.1 videos with smooth motion, preserved style, and strong scene coherence. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

google/gemini-2.5-flash-image-preview/text-to-image
text-to-image

google/gemini-2.5-flash-image-preview/text-to-image

Google Gemini 2.5 Flash Text-to-Image delivers state-of-the-art text-to-image generation and image editing with previews. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

Google Models

Google AI Models on WaveSpeedAI provide a comprehensive suite of generative AI tools for video, image, music, audio, and speech creation. The collection includes Veo for cinematic AI video generation, Imagen for high-quality image creation, Nano Banana and Nano Banana Lite for fast creative image workflows, Nano Banana Pro for premium visual production, Omni Flash for lightweight multimodal generation, Lyric for AI music creation, and Gemini Text-to-Speech for natural voice synthesis.

Built for creators, developers, marketers, and AI applications, Google AI Models support text-to-video, image-to-video, video extension, reference-based video generation, text-to-image, image editing, AI music generation, and text-to-speech workflows. These models combine strong prompt understanding, realistic motion, high visual fidelity, synchronized audio-video generation, fast iteration, and scalable API access for professional creative production.

Core Model Capabilities

Cinematic AI Video Generation:

Use Veo models to generate cinematic videos from text prompts, still images, start-end frames, or reference videos. Veo supports realistic motion, natural lighting, camera control, synchronized audio, and smooth scene continuity for storytelling, ads, product videos, and social media content.

Video Extension:

Extend existing Veo-generated videos into longer continuous clips while preserving motion style, framing, lighting, scene continuity, and synchronized audio. Fast variants are suitable for rapid previews, creative iteration, and multi-branch story continuation.

AI Image Generation:

Use Imagen, Nano Banana, Nano Banana Lite, Nano Banana Pro, Gemini, and Omni Flash models to create high-quality images from prompts for portraits, product visuals, key art, social media content, blog images, design concepts, and commercial assets.

AI Image Editing:

Transform existing images with prompt-guided editing, context-aware refinement, style adjustment, identity preservation, lighting consistency, and region-aware modifications.

Lightweight Creative Workflows:

Nano Banana, Nano Banana Lite, Gemini Flash, and Omni Flash provide faster and more cost-efficient generation options for everyday creative tasks, rapid prototyping, preview workflows, and high-volume image production.

Premium Image Workflows:

Imagen 4 Ultra, Nano Banana Pro, Nano Banana Pro Ultra, and Nano Banana Pro Multi support higher-fidelity image generation, improved prompt control, multi-reference consistency, and premium visual output for hero shots, advertising campaigns, brand assets, and professional creative projects.

AI Music Generation:

Lyric models generate high-quality music from prompts, supporting background scoring, social media content, soundtrack creation, and professional audio production workflows.

Text-to-Speech:

Gemini Text-to-Speech models provide natural and expressive voice synthesis for narration, dialogue, avatars, education, product explainers, and multilingual audio content.

Google AI Models on WaveSpeedAI give creators and developers fast access to Google's video, image, music, audio, and speech generation models with scalable APIs, flexible pricing, and production-ready creative capabilities.

Google Models API — 价格与性能

通过单一 REST API 运行 Google Models 系列中的任意模型。按生成计费 — 无订阅、无最低消费 — 在 99.9% 可用性的基础设施上提供行业领先的延迟。

为什么在 WaveSpeedAI 上运行 Google Models

透明定价

每个 Google Models 模型都有按调用计价。价格在每个模型的页面上列出 — 不收取额外的平台费。

为低延迟优化

大多数 Google Models 图像模型在 2 秒内完成。视频和 3D 模型比自托管方案快数倍。

99.9% 可用性

多区域故障转移和自动重试可确保您的生产流量保持在线 — 即使在供应商故障期间。

常见问题

Google Models API 多少钱?+

每个模型在其模型页面上都列有自己的按调用价格。我们按每次成功生成计费,没有订阅费或最低消费。

Google Models 模型在 WaveSpeedAI 上有多快?+

本系列中的图像模型通常在 2 秒内完成。视频和 3D 模型取决于时长和分辨率,但通常比自托管运行快数倍。

不用信用卡可以试用 API 吗?+

可以 — 每个账户在注册时获得 $1 的免费额度,足以在不使用信用卡的情况下试用大多数 Google Models 模型。

有速率限制吗?+

标准账户有充足的并发任务限制。企业版计划提供自定义 RPM、更高并发和专用容量 — 详情请联系销售。