Veo 3.1 API
Google Veo 3.1 支援帶同步原生音訊的 1080p 文字轉影片。提供 Standard、Fast、Lite 三種等級,包含文字轉影片、圖片轉影片、參考轉影片與 video-extend,Lite 等級另有 start-end-to-video。
Standard 等級提供 Veo 的完整畫質;Fast 等級用於迭代;Lite 等級用於預視覺化與大量作業。參考轉影片可接受參考圖片以維持身分一致;video-extend 以 7 秒為單位延伸,並輸出合併後的單一片段。
概覽
關於 Veo 3.1 API
Veo 3.1 的功能、它在 Google 模型陣容中的定位,以及團隊選用它的原因。
Veo 3.1 是來自 Google 的影片生成模型,可透過 WaveSpeedAI REST API 使用。Google Veo 3.1 支援帶同步原生音訊的 1080p 文字轉影片。提供 Standard、Fast、Lite 三種等級,包含文字轉影片、圖片轉影片、參考轉影片與 video-extend,Lite 等級另有 start-end-to-video。
Standard 等級提供 Veo 的完整畫質;Fast 等級用於迭代;Lite 等級用於預視覺化與大量作業。參考轉影片可接受參考圖片以維持身分一致;video-extend 以 7 秒為單位延伸,並輸出合併後的單一片段。
WaveSpeedAI 上的 Veo 3.1 系列提供 11 個 REST 端點,涵蓋 Text-To-Video, Video-Extend, Reference-To-Video, Image-To-Video 等工作流程。每個變體都有各自的定價、參數選項與範例輸出 — 請挑選符合您輸入模態與生產限制的版本,或使用同一組 API 金鑰呼叫多個變體,組合成多步驟的處理流程。
使用您呼叫其他 1,000+ 個 WaveSpeedAI AI 模型時相同的 API 金鑰、帳單帳戶與速率限制額度來執行 Veo 3.1。無需另外設定供應商、無需各家 SDK、無需處理各家不同的速率限制 — 一次整合,涵蓋從文字轉圖片、文字轉影片,到音訊合成、3D 生成、放大與編輯的所有功能。
端點
所有 Veo 3.1 API 端點
WaveSpeedAI 目前提供 11 個 Veo 3.1 端點 — 請選擇符合您工作流程的變體。
/filters:quality(82)/media/images/20260408104718_8uxbq2i4.webp)
Veo3.1 Lite Text To Video
Google Veo 3.1 Lite Text-to-Video generates high-fidelity 720p or 1080p videos with natively generated audio from text prompts. Lightweight variant optimized for cost efficiency. Supports landscape and portrait aspect ratios, dialogue with lip-sync, and customizable duration. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
/filters:quality(82)/media/images/20260408104740_drl8tu9l.webp)
Veo3.1 Video Extend
Extend and continue Veo 3.1 videos with smooth motion, preserved style, and strong scene coherence. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
/filters:quality(82)/media/images/1778749312530573096_makufpyI.webp)
Veo3.1 Fast Reference To Video
Google Veo 3.1 Fast Reference to Video is a fast AI reference-to-video generation model that creates 8-second videos from up to three reference images using the official Veo predictLongRunning endpoint with referenceImages assets. Ready-to-use REST inference API for product videos, character consistency, branded visual storytelling, social media clips, advertising creatives, and professional reference-based video generation workflows with simple integration, no coldstarts, and affordable pricing.
/filters:quality(82)/media/images/20260408104757_ct9o5p9o.webp)
Veo3.1 Fast Video Extend
Extend Veo 3.1 videos in 7-second steps with the Fast endpoint—quick, coherent continuation that preserves style and motion, output as a single merged clip. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
/filters:quality(82)/media/images/20260408104805_tr00kyv7.webp)
Veo3.1 Fast Text To Video
Google Veo 3.1 Fast creates text-to-video with native 1080p and synchronized audio, delivering high-quality videos for creators. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
/filters:quality(82)/media/images/20260408104748_nwaixdpw.webp)
Veo3.1 Reference To Video
Google Veo3.1 Reference-to-Video performs image-to-video generation that preserves a specific subject's appearance and identity from provided reference images, enabling consistent character or product motion across frames. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
/filters:quality(82)/media/images/20260408104714_ut4h5kku.webp)
Veo3.1 Lite Start End To Video
Google Veo 3.1 Lite Start-End-to-Video generates high-fidelity videos by interpolating between a start image and an optional end image. Supports 720p and 1080p resolutions, landscape and portrait aspect ratios, and native audio generation. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
/filters:quality(82)/media/images/20260408104752_1voe0l4b.webp)
Veo3.1 Text To Video
Google Veo 3.1 converts text prompts into videos with synchronized audio at native 1080p for high-quality outputs. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
/filters:quality(82)/media/images/20260408104708_dmnbefn0.webp)
Veo3.1 Lite Image To Video
Google Veo 3.1 Lite Image-to-Video transforms static images into high-fidelity 720p or 1080p videos with natively generated audio. Supports many interpolation use cases, landscape and portrait aspect ratios, and customizable duration. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
/filters:quality(82)/media/images/20260408104802_aaez2mmf.webp)
Veo3.1 Fast Image To Video
Google Veo 3.1 Fast is an Image-to-Video model with native 1080p output for high-detail videos from images and fast performance. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
/filters:quality(82)/media/images/20260408104744_xfvl07s5.webp)
Veo3.1 Image To Video
Google Veo 3.1 is an Image-to-Video model that converts images into high-quality videos with native 1080P output for enhanced detail and creative flexibility. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
範例
看看 Veo 3.1 的實際效果
由 Veo 3.1 API 生成的真實輸出。將滑鼠移到任一影片上即可預覽,點擊可開啟完整尺寸檢視器。
使用方式
如何使用 Veo 3.1 API
從註冊到完成生成,只要四個步驟。完整的 Python、Node.js 與 cURL 範例請見下方的 API 區段。
- 01
取得 API 金鑰
註冊 WaveSpeedAI 帳號,並從控制台複製您的 API 金鑰。新帳號會獲贈免費入門點數 — 足以在開始計費前執行數十次 Playground。
- 02
提交 prediction
將您的輸入以 JSON 格式 POST 到 https://api.wavespeed.ai/api/v3/google/veo3.1/text-to-video。端點會立即回傳 prediction id — 生成是非同步進行的,因此推論期間您不需要保持連線。
- 03
輪詢完成狀態
GET https://api.wavespeed.ai/api/v3/predictions/{request_id}/result。狀態為 completed 時回傳輸出;狀態為 failed、cancelled、timeout 或 deleted 時停止並回報錯誤;其他狀態則持續輪詢。
- 04
讀取輸出網址
當狀態為 "completed" 時,從 data.outputs[0] 讀取網址。該網址指向 WaveSpeedAI CDN 上您生成的媒體 — 依您呼叫的 Veo 3.1 變體,可能是圖片、影片、音訊或 3D 檔案。
應用情境
您可以用 Veo 3.1 打造什麼
開發者與創作者使用 Veo 3.1 API 的常見工作流程。
具原生 1080p 與音訊的文字轉影片
google/veo3.1/text-to-video 將文字提示詞轉為帶同步音訊、原生 1080p 的影片,無須放大處理即可交付。相同的提示詞格式適用於 Standard、Fast 與 Lite 各等級。
以參考轉影片維持身分一致
google/veo3.1/reference-to-video 在提供的參考圖片基礎上進行圖片轉影片,並保留特定主體的外觀。Fast 等級的參考轉影片使用 Veo 官方的 predictLongRunning 端點,最多支援三張參考圖片,輸出 8 秒影片。
分級定價:Standard / Fast / Lite
Standard 用於交付,Fast 用於迭代,Lite 用於預視覺化與大量作業。Lite 在畫質上有所取捨;只要更換端點 URL 即可切換等級,不必重寫提示詞。
以 7 秒為單位的 video-extend
google/veo3.1/video-extend 可延續現有的 Veo 影片,動態流暢並保留風格,輸出為單一合併片段,不需要另行拼接多個片段。Fast 等級的 video-extend(/extension)是反覆延伸時較便宜的選擇。
首尾幀插值(Lite 等級)
google/veo3.1-lite/start-end-to-video 會在起始圖片與選用的結束圖片之間插值,生成中間的連接動態。支援 720p 與 1080p、橫向與直向長寬比。適合在 Lite 價位下進行動態分鏡與關鍵幀驅動的工作流程。
原生 1080p 的圖片轉影片
google/veo3.1/image-to-video 將圖片轉為細節更豐富、創意空間更大的 1080p 影片,當您從一張關鍵靜態圖而非純文字簡報出發時,這是合適的選擇。
技巧
Veo 3.1 提示詞技巧
讓 Veo 3.1 產出更好結果的實用建議 — 整理自在實際生產流程中,影片模型通用的有效做法。
- 01
各等級皆為原生 1080p 輸出
目錄說明:「原生 1080p,輸出品質高」。所有 Veo 3.1 版本皆直接以 1080p 生成,交付前無須放大處理。若頻寬更重要,Lite 等級也支援 720p。
- 02
Lite 用於迭代,Fast 用於精修,Standard 用於交付
Standard、Fast 與 Lite 使用相同的提示詞格式,只需切換端點 URL,不必重寫。以 Lite 迭代提示詞方向,以 Fast 精修提升品質,需要最高保真度時再用 Standard 交付。
- 03
以參考轉影片維持身分一致
google/veo3.1/reference-to-video 可從參考圖片保留特定主體的外觀。Fast 等級的版本每次生成最多支援 3 張參考圖片。適用於主持人影片、品牌角色與系列化內容。
- 04
Video-extend 以 7 秒為單位
google/veo3.1-fast/video-extend 以 7 秒為單位延伸 Veo 3.1 影片,動態連貫且風格不變,輸出為單一合併片段,而不是需要在後製拼接的多個片段。
- 05
首尾幀插值僅限 Lite
google/veo3.1-lite/start-end-to-video 會在起始圖片與選用的結束圖片之間插值。支援 720p / 1080p、橫向與直向。僅限 Lite,Standard 與 Fast 不提供。
- 06
無聲影片請關閉音訊
Standard 文字轉影片端點提供 generate_audio 參數。B-roll、社群貼文,以及之後本來就要在後製替換音訊的影片,都可將它關閉。
定價
Veo 3.1 API 定價
依輸出計費。最終費用會依您在各變體 Playground 中設定的參數(解析度、時長、輸出數量、參考素材)而調整。
| 端點 | 類型 | 起始價格 |
|---|---|---|
| google/ | text-to-video | $0.30 |
| google/ | video-extend | $3.20 |
| google/ | reference-to-video | $0.64 |
| google/ | video-extend | $1.20 |
| google/ | text-to-video | $1.20 |
| google/ | reference-to-video | $3.20 |
| google/ | image-to-video | $0.40 |
| google/ | text-to-video | $3.20 |
| google/ | image-to-video | $0.30 |
| google/ | image-to-video | $1.20 |
| google/ | image-to-video | $3.20 |
API
呼叫 Veo 3.1 API
請到 wavespeed.ai/accesskey 註冊取得 API 金鑰,然後透過 REST 提交 prediction。Playground 會為任何輸入組合產生可直接貼上的範例。
POSThttps://api.wavespeed.ai/api/v3/google/veo3.1/text-to-video
# 1. Submit the prediction.
SUBMIT_RESPONSE=$(curl --silent --show-error --fail-with-body \
-X POST "https://api.wavespeed.ai/api/v3/google/veo3.1/text-to-video" \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $WAVESPEED_API_KEY" \
-d '{
"prompt": "A cinematic shot of a city at sunset, soft golden light",
"aspect_ratio": "16:9",
"duration": 8,
"resolution": "1080p",
"generate_audio": true
}')
TASK=$(printf '%s' "$SUBMIT_RESPONSE" | jq 'if has("data") then .data else . end')
PREDICTION_ID=$(printf '%s' "$TASK" | jq -r '.id')
if [ -z "$PREDICTION_ID" ] || [ "$PREDICTION_ID" = "null" ]; then
printf 'Submission response did not contain a prediction id
' >&2
exit 1
fi
RESULT_URL="https://api.wavespeed.ai/api/v3/predictions/$PREDICTION_ID/result"
# 2. Poll until the prediction finishes.
while true; do
RESPONSE=$(curl --silent --show-error --fail-with-body "$RESULT_URL" \
-H "Authorization: Bearer $WAVESPEED_API_KEY")
RESULT=$(printf '%s' "$RESPONSE" | jq 'if has("data") then .data else . end')
STATUS=$(printf '%s' "$RESULT" | jq -r '.status')
case "$STATUS" in
completed) printf '%s\n' "$RESULT" | jq '.outputs'; break ;;
failed|cancelled|timeout|deleted) printf '%s\n' "$RESULT" | jq . >&2; exit 1 ;;
*) sleep 2 ;;
esac
doneconst submitUrl = "https://api.wavespeed.ai/api/v3/google/veo3.1/text-to-video";
const apiKey = process.env.WAVESPEED_API_KEY;
if (!apiKey) throw new Error('Set WAVESPEED_API_KEY');
async function requestJson(url, options = {}) {
const response = await fetch(url, options);
if (!response.ok) throw new Error(await response.text());
return response.json();
}
// 1. Submit the prediction.
const body = await requestJson(submitUrl, {
method: "POST",
headers: {
"Authorization": `Bearer ${apiKey}`,
"Content-Type": "application/json",
},
body: JSON.stringify({
"prompt": "A cinematic shot of a city at sunset, soft golden light",
"aspect_ratio": "16:9",
"duration": 8,
"resolution": "1080p",
"generate_audio": true
}),
});
const task = body.data ?? body;
const resultUrl = `https://api.wavespeed.ai/api/v3/predictions/${task.id}/result`;
// 2. Poll until the prediction finishes.
while (true) {
const resultBody = await requestJson(resultUrl, {
headers: { "Authorization": `Bearer ${apiKey}` },
});
const result = resultBody.data ?? resultBody;
if (result.status === "completed") {
console.log(result.outputs);
break;
}
if (["failed", "cancelled", "timeout", "deleted"].includes(result.status)) throw new Error(JSON.stringify(result));
await new Promise(resolve => setTimeout(resolve, 2000));
}import json
import os
import time
from urllib.request import Request, urlopen
api_key = os.environ["WAVESPEED_API_KEY"]
headers = {"Authorization": f"Bearer {api_key}", "Content-Type": "application/json"}
payload = {
"prompt": "A cinematic shot of a city at sunset, soft golden light",
"aspect_ratio": "16:9",
"duration": 8,
"resolution": "1080p",
"generate_audio": True
}
def request_json(url, data=None):
request = Request(url, data=data, headers=headers, method="POST" if data else "GET")
with urlopen(request) as response:
return json.load(response)
# 1. Submit the prediction.
body = request_json("https://api.wavespeed.ai/api/v3/google/veo3.1/text-to-video", json.dumps(payload).encode())
task = body.get("data", body)
result_url = f"https://api.wavespeed.ai/api/v3/predictions/{task['id']}/result"
# 2. Poll until the prediction finishes.
while True:
result_body = request_json(result_url)
result = result_body.get("data", result_body)
status = result.get("status")
if status == "completed":
print(result.get("outputs", []))
break
if status in {"failed", "cancelled", "timeout", "deleted"}:
raise RuntimeError(result)
time.sleep(2)比較
Veo 3.1 與其他選擇比較
在 WaveSpeedAI 上,何時該選擇 Veo 3.1 而非類似的模型。
Veo 3.1 vs Seedance 2.0
Seedance 2.0 在所有等級皆具原生音訊,並有快速的 Turbo 等級。Veo 3.1 在最高等級價格較高,但 Lite 等級在成本上具有競爭力,且 Veo 在人臉的寫實度上口碑更佳。
Veo 3.1 vs Kling 3.0
Kling 3.0 提供 Pro 與 4K 等級以及動作控制端點。Veo 3.1 則有三種價格等級(Standard / Fast / Lite),並將參考轉影片與首尾幀插值列為核心端點,功能組合各有不同。
Veo 3.1 vs Wan 2.7
Wan 2.7 在同系列中提供參考轉影片、圖片編輯與文字轉圖片版本,跨模態工具組更廣。Veo 3.1 的 Standard 等級涵蓋分級的成本最佳化(Lite 至 Standard),以及 Wan 以不同方式處理的 7 秒 video-extend。
常見問題
Veo 3.1 API — 常見問題
定價、授權、整合 — 關於在 WaveSpeedAI 上執行 Veo 3.1 的常見問題。
什麼是 Veo 3.1 API?
Veo 3.1 是 Google 的影片生成模型,在 WaveSpeedAI 上以 REST API 提供。Google Veo 3.1 支援帶同步原生音訊的 1080p 文字轉影片。提供 Standard、Fast、Lite 三種等級,包含文字轉影片、圖片轉影片、參考轉影片與 video-extend,Lite 等級另有 start-end-to-video。您可以透過程式呼叫,也可以在上方連結的 Playground 試用。
如何呼叫 Veo 3.1 API?
註冊 WaveSpeedAI 帳號,從 /accesskey 複製您的 API 金鑰,然後將您的輸入以 JSON 格式 POST 到 https://api.wavespeed.ai/api/v3/google/veo3.1/text-to-video。端點會回傳 prediction id。請從約每 2 秒輪詢一次結果端點開始,長時間任務可拉長間隔,並在任何終止狀態時停止。上方提供適用於生產環境的 Python / Node.js / cURL 範例。
Veo 3.1 API 的費用是多少?
Veo 3.1 每次執行 $0.30 起。實際費用會依您設定的參數(解析度、時長、輸出數量、參考素材)而調整。Playground 中「生成」按鈕旁的即時費用預覽會顯示您目前輸入的確切價格。
有哪些 Veo 3.1 變體可用?
WaveSpeedAI 提供 11 個已上線的 Veo 3.1 端點:google/veo3.1-lite/text-to-video, google/veo3.1/video-extend, google/veo3.1-fast/reference-to-video, google/veo3.1-fast/video-extend, google/veo3.1-fast/text-to-video, google/veo3.1/reference-to-video, google/veo3.1-lite/start-end-to-video, google/veo3.1/text-to-video等。每個變體都有各自的 Playground 頁面與定價。
Veo 3.1 的輸出可以商用嗎?
商業使用權依 Google 的模型授權而定。多數 Google 模型允許商用輸出;請參閱各模型 Playground 頁面中的授權摘要,以及 WaveSpeedAI 的服務條款了解平台層級的條件。
為什麼要在 WaveSpeedAI 上使用 Veo 3.1,而不是直接使用?
一組 API 金鑰、一個帳單帳戶,即可使用 Veo 3.1 以及來自其他供應商的 1,000+ 個 AI 模型。無需設定各家 SDK、無需處理各自獨立的速率限制、無需為每個供應商重寫整合程式碼。價格通常與 Google 的直接 API 相當或更低。
提供者
關於 Google
Veo 3.1 背後的團隊,以及 Google 在 WaveSpeedAI 上的完整模型陣容。
Google 的 AI 研究主要在 Google DeepMind 與 Google Research 進行。其圖像與影片模型,包括 Imagen、Veo,以及 Nano Banana(Gemini 3 Image)等 Gemini 系列多模態模型,與整個 Gemini 產品線共用架構與訓練基礎設施。輸出以準確的文字渲染、廣泛的風格涵蓋與商用等級授權見長。
在 WaveSpeedAI 上開始使用 Veo 3.1 打造應用
註冊即贈免費入門點數。一組 API 金鑰,涵蓋來自 Google 與所有其他供應商的 1,000+ 個 AI 模型。