Veo 3.1 API
Google Veo 3.1——文生视频,带 1080p 同步原生音频。提供 Standard、Fast、Lite 三个档位,包含文生视频、图生视频、参考生视频和视频延长,Lite 档位另有首尾帧生视频。
Standard 档位提供 Veo 的完整画质;Fast 档位用于迭代;Lite 档位适合预演和大批量工作。参考生视频接受参考图像,用于保持身份的生成;视频延长以 7 秒为步长,输出合并后的单个片段。
概览
关于 Veo 3.1 API
Veo 3.1 能做什么、它在 Google 模型阵容中的定位,以及团队选择它的原因。
Veo 3.1 是 Google 推出的视频生成模型,可通过 WaveSpeedAI REST API 使用。Google Veo 3.1——文生视频,带 1080p 同步原生音频。提供 Standard、Fast、Lite 三个档位,包含文生视频、图生视频、参考生视频和视频延长,Lite 档位另有首尾帧生视频。
Standard 档位提供 Veo 的完整画质;Fast 档位用于迭代;Lite 档位适合预演和大批量工作。参考生视频接受参考图像,用于保持身份的生成;视频延长以 7 秒为步长,输出合并后的单个片段。
WaveSpeedAI 上的 Veo 3.1 系列提供 11 个 REST 端点,涵盖 Text-To-Video, Video-Extend, Reference-To-Video, Image-To-Video 个工作流。每个变体都有各自的定价、参数选项和示例输出——请选择与你的输入模态和生产约束相匹配的那一个,或使用同一个 API 密钥调用多个变体,组合成多步骤流水线。
使用与 WaveSpeedAI 上其他 1,000 多个 AI 模型相同的 API 密钥、账单账户和速率限制来运行 Veo 3.1。无需单独对接供应商,无需各家 SDK,也无需应对各家不同的速率限制——一次集成即可覆盖从文生图、文生视频到音频合成、3D 生成、放大和编辑的全部能力。
端点
全部 Veo 3.1 API 端点
WaveSpeedAI 现已提供 11 个 Veo 3.1 端点——请选择适合你工作流的变体。
/filters:quality(82)/media/images/20260408104718_8uxbq2i4.webp)
Veo3.1 Lite Text To Video
Google Veo 3.1 Lite Text-to-Video generates high-fidelity 720p or 1080p videos with natively generated audio from text prompts. Lightweight variant optimized for cost efficiency. Supports landscape and portrait aspect ratios, dialogue with lip-sync, and customizable duration. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
/filters:quality(82)/media/images/20260408104740_drl8tu9l.webp)
Veo3.1 Video Extend
Extend and continue Veo 3.1 videos with smooth motion, preserved style, and strong scene coherence. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
/filters:quality(82)/media/images/1778749312530573096_makufpyI.webp)
Veo3.1 Fast Reference To Video
Google Veo 3.1 Fast Reference to Video is a fast AI reference-to-video generation model that creates 8-second videos from up to three reference images using the official Veo predictLongRunning endpoint with referenceImages assets. Ready-to-use REST inference API for product videos, character consistency, branded visual storytelling, social media clips, advertising creatives, and professional reference-based video generation workflows with simple integration, no coldstarts, and affordable pricing.
/filters:quality(82)/media/images/20260408104757_ct9o5p9o.webp)
Veo3.1 Fast Video Extend
Extend Veo 3.1 videos in 7-second steps with the Fast endpoint—quick, coherent continuation that preserves style and motion, output as a single merged clip. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
/filters:quality(82)/media/images/20260408104805_tr00kyv7.webp)
Veo3.1 Fast Text To Video
Google Veo 3.1 Fast creates text-to-video with native 1080p and synchronized audio, delivering high-quality videos for creators. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
/filters:quality(82)/media/images/20260408104748_nwaixdpw.webp)
Veo3.1 Reference To Video
Google Veo3.1 Reference-to-Video performs image-to-video generation that preserves a specific subject's appearance and identity from provided reference images, enabling consistent character or product motion across frames. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
/filters:quality(82)/media/images/20260408104714_ut4h5kku.webp)
Veo3.1 Lite Start End To Video
Google Veo 3.1 Lite Start-End-to-Video generates high-fidelity videos by interpolating between a start image and an optional end image. Supports 720p and 1080p resolutions, landscape and portrait aspect ratios, and native audio generation. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
/filters:quality(82)/media/images/20260408104752_1voe0l4b.webp)
Veo3.1 Text To Video
Google Veo 3.1 converts text prompts into videos with synchronized audio at native 1080p for high-quality outputs. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
/filters:quality(82)/media/images/20260408104708_dmnbefn0.webp)
Veo3.1 Lite Image To Video
Google Veo 3.1 Lite Image-to-Video transforms static images into high-fidelity 720p or 1080p videos with natively generated audio. Supports many interpolation use cases, landscape and portrait aspect ratios, and customizable duration. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
/filters:quality(82)/media/images/20260408104802_aaez2mmf.webp)
Veo3.1 Fast Image To Video
Google Veo 3.1 Fast is an Image-to-Video model with native 1080p output for high-detail videos from images and fast performance. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
/filters:quality(82)/media/images/20260408104744_xfvl07s5.webp)
Veo3.1 Image To Video
Google Veo 3.1 is an Image-to-Video model that converts images into high-quality videos with native 1080P output for enhanced detail and creative flexibility. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
示例
看看 Veo 3.1 的实际效果
由 Veo 3.1 API 生成的真实输出。悬停在任意视频上即可预览,点击可打开全尺寸查看器。
使用方法
如何使用 Veo 3.1 API
从注册到完成一次生成,只需四步。完整的 Python、Node.js 和 cURL 示例见下方的 API 部分。
- 01
获取 API 密钥
注册 WaveSpeedAI 账号,并从控制台复制你的 API 密钥。新账号附带免费体验额度——足够在开始计费前把 Playground 运行几十次。
- 02
提交预测
把你的输入以 JSON 形式 POST 到 https://api.wavespeed.ai/api/v3/google/veo3.1/text-to-video。端点会立即返回预测 ID——生成是异步的,因此推理期间你无需保持连接。
- 03
轮询完成状态
GET https://api.wavespeed.ai/api/v3/predictions/{request_id}/result。状态为 completed 时返回输出;状态为 failed、cancelled、timeout 或 deleted 时以错误终止;其他任何状态都继续轮询。
- 04
读取输出 URL
状态变为 "completed" 后,从 data.outputs[0] 读取 URL。该 URL 指向 WaveSpeedAI CDN 上你生成的媒体——具体是图片、视频、音频还是 3D 文件,取决于你调用的 Veo 3.1 变体。
应用场景
用 Veo 3.1 可以构建什么
开发者和创作者使用 Veo 3.1 API 的常见工作流。
文生视频,原生 1080p 音频
google/veo3.1/text-to-video 将文本提示词转换为带同步音频的原生 1080p 视频——无需放大处理即可达到可交付的分辨率。同样的提示词格式适用于 Standard、Fast 和 Lite 档位。
参考生视频,保持身份一致
google/veo3.1/reference-to-video 在图生视频的同时,保留所提供参考图像中特定主体的外观。Fast 档位的参考生视频使用 Veo 官方的 predictLongRunning 端点,最多支持三张参考图像,输出 8 秒视频。
分档定价:Standard / Fast / Lite
Standard 用于交付,Fast 用于迭代,Lite 用于预演和大批量工作。Lite 在画质上有所取舍;通过端点 URL 切换档位即可,无需重写提示词。
以 7 秒为步长的视频延长
google/veo3.1/video-extend 以流畅的运动和一致的风格续写现有 Veo 片段——输出是合并后的单个片段,而不是需要自行拼接的多个片段。Fast 档位的视频延长(/extension)是迭代延长时更便宜的选择。
首尾帧插值(Lite 档位)
google/veo3.1-lite/start-end-to-video 在起始图像和可选的结束图像之间插值,生成衔接的运动。支持 720p 和 1080p,横屏和竖屏画幅。适合在 Lite 价格档位上进行动态分镜和关键帧驱动的工作流。
原生 1080p 图生视频
google/veo3.1/image-to-video 将图像转换为 1080p 视频,细节更丰富、创作更灵活——从关键静态图而非纯文本说明起步时,这是合适的选择。
技巧
Veo 3.1 提示词技巧
让 Veo 3.1 输出更好结果的实用建议——总结自生产流水线中各类视频模型都适用的做法。
- 01
各档位均为原生 1080p 输出
目录宣称:「高质量输出为原生 1080p」。所有 Veo 3.1 变体都直接生成 1080p——交付前无需放大处理。当带宽更重要时,Lite 档位也支持 720p。
- 02
用 Lite 迭代,用 Fast 精修,用 Standard 交付
Standard、Fast 和 Lite 使用相同的提示词格式——切换端点 URL 即可,无需重写。用 Lite 迭代提示词方向,用 Fast 精修以获得更高画质,需要最高保真度时用 Standard 交付。
- 03
参考生视频,保持身份一致
google/veo3.1/reference-to-video 会保留参考图像中特定主体的外观。Fast 档位的版本每次生成最多支持 3 张参考图像。适合主持人视频、品牌角色和系列化内容。
- 04
视频延长以 7 秒为步长
google/veo3.1-fast/video-extend 以 7 秒为步长延长 Veo 3.1 视频,运动连贯、风格一致——输出是合并后的单个片段,而不是需要在后期拼接的多个片段。
- 05
首尾帧插值仅限 Lite
google/veo3.1-lite/start-end-to-video 在起始图像和可选的结束图像之间插值。支持 720p / 1080p,横屏和竖屏。仅限 Lite——Standard 和 Fast 不提供。
- 06
无声素材可关闭音频
Standard 文生视频端点开放了 generate_audio 参数。B-roll、社交帖子以及之后本来就要替换音频的片段,可将其关闭。
定价
Veo 3.1 API 定价
按输出计费。最终费用会随你在各变体 Playground 中设置的参数(分辨率、时长、输出数量、参考素材)而变化。
| 端点 | 类型 | 起步价 |
|---|---|---|
| google/ | text-to-video | $0.30 |
| google/ | video-extend | $3.20 |
| google/ | reference-to-video | $0.64 |
| google/ | video-extend | $1.20 |
| google/ | text-to-video | $1.20 |
| google/ | reference-to-video | $3.20 |
| google/ | image-to-video | $0.40 |
| google/ | text-to-video | $3.20 |
| google/ | image-to-video | $0.30 |
| google/ | image-to-video | $1.20 |
| google/ | image-to-video | $3.20 |
API
调用 Veo 3.1 API
在 wavespeed.ai/accesskey 注册并获取 API 密钥,然后通过 REST 提交预测。Playground 可以为任意输入组合生成可直接粘贴的示例代码。
POSThttps://api.wavespeed.ai/api/v3/google/veo3.1/text-to-video
# 1. Submit the prediction.
SUBMIT_RESPONSE=$(curl --silent --show-error --fail-with-body \
-X POST "https://api.wavespeed.ai/api/v3/google/veo3.1/text-to-video" \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $WAVESPEED_API_KEY" \
-d '{
"prompt": "A cinematic shot of a city at sunset, soft golden light",
"aspect_ratio": "16:9",
"duration": 8,
"resolution": "1080p",
"generate_audio": true
}')
TASK=$(printf '%s' "$SUBMIT_RESPONSE" | jq 'if has("data") then .data else . end')
PREDICTION_ID=$(printf '%s' "$TASK" | jq -r '.id')
if [ -z "$PREDICTION_ID" ] || [ "$PREDICTION_ID" = "null" ]; then
printf 'Submission response did not contain a prediction id
' >&2
exit 1
fi
RESULT_URL="https://api.wavespeed.ai/api/v3/predictions/$PREDICTION_ID/result"
# 2. Poll until the prediction finishes.
while true; do
RESPONSE=$(curl --silent --show-error --fail-with-body "$RESULT_URL" \
-H "Authorization: Bearer $WAVESPEED_API_KEY")
RESULT=$(printf '%s' "$RESPONSE" | jq 'if has("data") then .data else . end')
STATUS=$(printf '%s' "$RESULT" | jq -r '.status')
case "$STATUS" in
completed) printf '%s\n' "$RESULT" | jq '.outputs'; break ;;
failed|cancelled|timeout|deleted) printf '%s\n' "$RESULT" | jq . >&2; exit 1 ;;
*) sleep 2 ;;
esac
doneconst submitUrl = "https://api.wavespeed.ai/api/v3/google/veo3.1/text-to-video";
const apiKey = process.env.WAVESPEED_API_KEY;
if (!apiKey) throw new Error('Set WAVESPEED_API_KEY');
async function requestJson(url, options = {}) {
const response = await fetch(url, options);
if (!response.ok) throw new Error(await response.text());
return response.json();
}
// 1. Submit the prediction.
const body = await requestJson(submitUrl, {
method: "POST",
headers: {
"Authorization": `Bearer ${apiKey}`,
"Content-Type": "application/json",
},
body: JSON.stringify({
"prompt": "A cinematic shot of a city at sunset, soft golden light",
"aspect_ratio": "16:9",
"duration": 8,
"resolution": "1080p",
"generate_audio": true
}),
});
const task = body.data ?? body;
const resultUrl = `https://api.wavespeed.ai/api/v3/predictions/${task.id}/result`;
// 2. Poll until the prediction finishes.
while (true) {
const resultBody = await requestJson(resultUrl, {
headers: { "Authorization": `Bearer ${apiKey}` },
});
const result = resultBody.data ?? resultBody;
if (result.status === "completed") {
console.log(result.outputs);
break;
}
if (["failed", "cancelled", "timeout", "deleted"].includes(result.status)) throw new Error(JSON.stringify(result));
await new Promise(resolve => setTimeout(resolve, 2000));
}import json
import os
import time
from urllib.request import Request, urlopen
api_key = os.environ["WAVESPEED_API_KEY"]
headers = {"Authorization": f"Bearer {api_key}", "Content-Type": "application/json"}
payload = {
"prompt": "A cinematic shot of a city at sunset, soft golden light",
"aspect_ratio": "16:9",
"duration": 8,
"resolution": "1080p",
"generate_audio": True
}
def request_json(url, data=None):
request = Request(url, data=data, headers=headers, method="POST" if data else "GET")
with urlopen(request) as response:
return json.load(response)
# 1. Submit the prediction.
body = request_json("https://api.wavespeed.ai/api/v3/google/veo3.1/text-to-video", json.dumps(payload).encode())
task = body.get("data", body)
result_url = f"https://api.wavespeed.ai/api/v3/predictions/{task['id']}/result"
# 2. Poll until the prediction finishes.
while True:
result_body = request_json(result_url)
result = result_body.get("data", result_body)
status = result.get("status")
if status == "completed":
print(result.get("outputs", []))
break
if status in {"failed", "cancelled", "timeout", "deleted"}:
raise RuntimeError(result)
time.sleep(2)对比
Veo 3.1 与其他方案对比
在 WaveSpeedAI 上,何时应选择 Veo 3.1 而不是同类模型。
Veo 3.1 对比 Seedance 2.0
Seedance 2.0 每个档位都带原生音频,并有快速的 Turbo 档位。Veo 3.1 的顶级档位更贵,但 Lite 档位在成本上颇具竞争力,且 Veo 在人脸写实方面的口碑更强。
Veo 3.1 对比 Kling 3.0
Kling 3.0 有 Pro 和 4K 档位及运动控制端点。Veo 3.1 提供三个价格档位(Standard / Fast / Lite),并把参考生视频和首尾帧插值作为一等端点——功能范围不同,侧重点也不同。
Veo 3.1 对比 Wan 2.7
Wan 2.7 在一个系列中提供参考生视频、图像编辑和文生图变体——跨模态工具集更广。Veo 3.1 的 Standard 档位涵盖分档的成本优化(从 Lite 到 Standard),以及 7 秒视频延长,而 Wan 对此的处理方式不同。
常见问题
Veo 3.1 API — 常见问题
定价、许可、集成——关于在 WaveSpeedAI 上运行 Veo 3.1 的常见问题。
Veo 3.1 API 是什么?
Veo 3.1 是 Google 的视频生成模型,在 WaveSpeedAI 上以 REST API 形式提供。Google Veo 3.1——文生视频,带 1080p 同步原生音频。提供 Standard、Fast、Lite 三个档位,包含文生视频、图生视频、参考生视频和视频延长,Lite 档位另有首尾帧生视频。你可以通过编程方式调用它,也可以在上方链接的 Playground 中试用。
如何调用 Veo 3.1 API?
注册 WaveSpeedAI 账号,从 /accesskey 复制你的 API 密钥,然后把输入以 JSON 形式 POST 到 https://api.wavespeed.ai/api/v3/google/veo3.1/text-to-video。端点会返回预测 ID。从大约每 2 秒一次开始轮询结果端点,长耗时任务可适当拉长间隔,并在任何终止状态时停止。上方有面向生产环境的 Python / Node.js / cURL 示例。
Veo 3.1 API 的费用是多少?
Veo 3.1 每次调用 $0.30 起。实际费用会随你设置的参数(分辨率、时长、输出数量、参考素材)而变化。Playground 中“生成”按钮旁的实时费用预览会显示你当前输入对应的准确价格。
有哪些 Veo 3.1 变体可用?
WaveSpeedAI 托管了 11 个已上线的 Veo 3.1 端点:google/veo3.1-lite/text-to-video, google/veo3.1/video-extend, google/veo3.1-fast/reference-to-video, google/veo3.1-fast/video-extend, google/veo3.1-fast/text-to-video, google/veo3.1/reference-to-video, google/veo3.1-lite/start-end-to-video, google/veo3.1/text-to-video等。每个变体都有自己的 Playground 页面和定价。
Veo 3.1 的输出可以商用吗?
商用权利遵循 Google 的模型许可。大多数 Google 模型允许商用输出;具体许可摘要请查看各模型的 Playground 页面,平台层面的条件请参阅 WaveSpeedAI 的服务条款。
为什么要在 WaveSpeedAI 上使用 Veo 3.1,而不是直接对接?
一个 API 密钥、一个账单账户,即可使用 Veo 3.1 以及来自其他提供方的 1,000 多个 AI 模型。无需逐家配置 SDK,无需应对各自独立的速率限制,也无需为每家重写集成代码。价格通常与 Google 直接提供的 API 持平或更低。
提供方
关于 Google
Veo 3.1 及 WaveSpeedAI 上 Google 更多模型背后的团队。
Google 的 AI 工作主要在 Google DeepMind 和 Google Research 进行。其图像和视频模型——Imagen、Veo,以及 Nano Banana(Gemini 3 Image)等 Gemini 系列多模态模型——与更广泛的 Gemini 产品线共享架构和训练基础设施。其输出以准确的文字渲染、广泛的风格覆盖和商用级许可著称。
在 WaveSpeedAI 上用 Veo 3.1 开始构建
注册即送免费体验额度。一个 API 密钥,即可使用来自 Google 及其他所有提供方的 1,000 多个 AI 模型。