Veo 3.1 API
Google の Veo 3.1 は、1080pの同期ネイティブオーディオ付きのテキストから動画に対応しています。3つの階層(Standard、Fast、Lite)に、テキストから動画、画像から動画、リファレンスから動画、video-extend があり、Lite 階層では start-end-to-video も利用できます。
Standard 階層は Veo のフル品質、Fast 階層は試行用、Lite 階層はプリビズや大量処理向けです。リファレンスから動画は参照画像でアイデンティティを保った生成ができ、video-extend は7秒単位で延長して、1本に統合されたクリップを出力します。
概要
Veo 3.1 APIについて
Veo 3.1でできること、Googleのモデルラインナップにおける位置づけ、そして多くのチームに選ばれる理由。
Veo 3.1はGoogleの動画生成モデルで、WaveSpeedAIのREST APIから利用できます。Google の Veo 3.1 は、1080pの同期ネイティブオーディオ付きのテキストから動画に対応しています。3つの階層(Standard、Fast、Lite)に、テキストから動画、画像から動画、リファレンスから動画、video-extend があり、Lite 階層では start-end-to-video も利用できます。
Standard 階層は Veo のフル品質、Fast 階層は試行用、Lite 階層はプリビズや大量処理向けです。リファレンスから動画は参照画像でアイデンティティを保った生成ができ、video-extend は7秒単位で延長して、1本に統合されたクリップを出力します。
WaveSpeedAIのVeo 3.1ファミリーには、Text-To-Video, Video-Extend, Reference-To-Video, Image-To-Videoのワークフローをカバーする11個のRESTエンドポイントがあります。各バリアントには、それぞれ独自の料金、パラメーター、作例があります。入力の種類と本番環境の制約に合うものを選ぶか、同じAPIキーで複数を呼び出して、多段階のパイプラインを組み立ててください。
WaveSpeedAIの他の1,000以上のAIモデルと同じAPIキー、請求アカウント、レート制限の枠組みでVeo 3.1を実行できます。ベンダーごとの設定も、プロバイダーごとのSDKも、ベンダーごとのレート制限も不要です。1つの連携で、テキストから画像、テキストから動画から、音声合成、3D生成、高画質化、編集まで、すべてカバーできます。
エンドポイント
Veo 3.1のAPIエンドポイント一覧
WaveSpeedAIで現在利用可能なVeo 3.1のエンドポイントは11個です。ワークフローに合ったバリアントを選んでください。
/filters:quality(82)/media/images/20260408104718_8uxbq2i4.webp)
Veo3.1 Lite Text To Video
Google Veo 3.1 Lite Text-to-Video generates high-fidelity 720p or 1080p videos with natively generated audio from text prompts. Lightweight variant optimized for cost efficiency. Supports landscape and portrait aspect ratios, dialogue with lip-sync, and customizable duration. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
/filters:quality(82)/media/images/20260408104740_drl8tu9l.webp)
Veo3.1 Video Extend
Extend and continue Veo 3.1 videos with smooth motion, preserved style, and strong scene coherence. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
/filters:quality(82)/media/images/1778749312530573096_makufpyI.webp)
Veo3.1 Fast Reference To Video
Google Veo 3.1 Fast Reference to Video is a fast AI reference-to-video generation model that creates 8-second videos from up to three reference images using the official Veo predictLongRunning endpoint with referenceImages assets. Ready-to-use REST inference API for product videos, character consistency, branded visual storytelling, social media clips, advertising creatives, and professional reference-based video generation workflows with simple integration, no coldstarts, and affordable pricing.
/filters:quality(82)/media/images/20260408104757_ct9o5p9o.webp)
Veo3.1 Fast Video Extend
Extend Veo 3.1 videos in 7-second steps with the Fast endpoint—quick, coherent continuation that preserves style and motion, output as a single merged clip. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
/filters:quality(82)/media/images/20260408104805_tr00kyv7.webp)
Veo3.1 Fast Text To Video
Google Veo 3.1 Fast creates text-to-video with native 1080p and synchronized audio, delivering high-quality videos for creators. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
/filters:quality(82)/media/images/20260408104748_nwaixdpw.webp)
Veo3.1 Reference To Video
Google Veo3.1 Reference-to-Video performs image-to-video generation that preserves a specific subject's appearance and identity from provided reference images, enabling consistent character or product motion across frames. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
/filters:quality(82)/media/images/20260408104714_ut4h5kku.webp)
Veo3.1 Lite Start End To Video
Google Veo 3.1 Lite Start-End-to-Video generates high-fidelity videos by interpolating between a start image and an optional end image. Supports 720p and 1080p resolutions, landscape and portrait aspect ratios, and native audio generation. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
/filters:quality(82)/media/images/20260408104752_1voe0l4b.webp)
Veo3.1 Text To Video
Google Veo 3.1 converts text prompts into videos with synchronized audio at native 1080p for high-quality outputs. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
/filters:quality(82)/media/images/20260408104708_dmnbefn0.webp)
Veo3.1 Lite Image To Video
Google Veo 3.1 Lite Image-to-Video transforms static images into high-fidelity 720p or 1080p videos with natively generated audio. Supports many interpolation use cases, landscape and portrait aspect ratios, and customizable duration. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
/filters:quality(82)/media/images/20260408104802_aaez2mmf.webp)
Veo3.1 Fast Image To Video
Google Veo 3.1 Fast is an Image-to-Video model with native 1080p output for high-detail videos from images and fast performance. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
/filters:quality(82)/media/images/20260408104744_xfvl07s5.webp)
Veo3.1 Image To Video
Google Veo 3.1 is an Image-to-Video model that converts images into high-quality videos with native 1080P output for enhanced detail and creative flexibility. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
作例
Veo 3.1の実力を見る
Veo 3.1 APIで実際に生成された出力です。動画にカーソルを合わせるとプレビュー、クリックすると原寸のビューアで開きます。
使い方
Veo 3.1 APIの使い方
登録から生成完了まで4ステップ。Python、Node.js、cURLの完全なサンプルは、下のAPIセクションにあります。
- 01
APIキーを取得
WaveSpeedAIのアカウントに登録し、ダッシュボードからAPIキーをコピーします。新規アカウントには無料のスタータークレジットが付くので、課金が始まる前にプレイグラウンドを数十回実行できます。
- 02
予測を送信
入力をJSONとしてhttps://api.wavespeed.ai/api/v3/google/veo3.1/text-to-videoにPOSTします。エンドポイントはすぐに予測IDを返します。生成は非同期なので、推論中に接続を開いたままにする必要はありません。
- 03
完了までポーリング
https://api.wavespeed.ai/api/v3/predictions/{request_id}/resultにGETします。completedなら出力を返し、failed、cancelled、timeout、deletedならエラーで停止し、それ以外のステータスの間はポーリングを続けます。
- 04
出力URLを読み取る
ステータスが"completed"になったら、data.outputs[0]からURLを読み取ります。URLは、WaveSpeedAIのCDN上にある生成されたメディアを指します。呼び出したVeo 3.1のバリアントに応じて、画像、動画、音声、3Dファイルのいずれかです。
活用例
Veo 3.1で作れるもの
開発者やクリエイターがVeo 3.1 APIでよく使うワークフロー。
ネイティブ1080pオーディオ付きのテキストから動画
google/veo3.1/text-to-video は、テキストプロンプトを、ネイティブ1080pの同期オーディオ付き動画に変換します。アップスケールの工程なしで、納品に使える解像度が得られます。同じプロンプト形式が Standard、Fast、Lite の各階層で使えます。
アイデンティティを保つリファレンスから動画
google/veo3.1/reference-to-video は、提供された参照画像から特定の被写体の見た目を保ったまま、画像から動画の生成を行います。Fast 階層のリファレンスから動画は、Veo 公式の predictLongRunning エンドポイントを使い、最大3枚の参照画像で8秒の出力に対応します。
階層型料金: Standard / Fast / Lite
納品には Standard、試行には Fast、プリビズや大量処理には Lite を使います。Lite は品質面でトレードオフがあります。階層は、プロンプトを書き直さずにエンドポイントURLで切り替えられます。
7秒単位の video-extend
google/veo3.1/video-extend は、既存の Veo クリップを、滑らかなモーションとスタイルを保ったまま続けます。出力は、つなぎ合わせる別々のセグメントではなく、1本に統合されたクリップです。繰り返し延長するなら、Fast 階層の video-extend(/extension)が安価な選択肢です。
開始・終了フレームの補間(Lite 階層)
google/veo3.1-lite/start-end-to-video は、開始画像と任意の終了画像の間を補間し、つなぐモーションを生成します。720pと1080p、横向きと縦向きのアスペクト比に対応。Lite の価格帯で、アニマティックやキーフレーム主体のワークフローに便利です。
ネイティブ1080pの画像から動画
google/veo3.1/image-to-video は、画像を、ディテールと創造の自由度が高まった1080pの動画に変換します。テキストだけのブリーフではなく、キーとなる静止画から始める場合に最適です。
コツ
Veo 3.1のプロンプトのコツ
Veo 3.1からより良い出力を得るための実践的なアドバイス。本番のパイプラインで動画モデル全般に通用するパターンに基づいています。
- 01
全階層でネイティブ1080p出力
カタログ上、「高品質な出力のためのネイティブ1080p」に対応しています。Veo 3.1 のすべてのバリアントは、1080pで直接生成するため、納品前のアップスケールの工程は不要です。帯域が重要な場合は、Lite 階層で720pも利用できます。
- 02
試行は Lite、仕上げは Fast、納品は Standard
同じプロンプト形式が、Standard、Fast、Lite のすべてで使えます。書き直さずに、エンドポイントのURLを切り替えるだけです。プロンプトの方向性は Lite で試し、より高い品質を求めて Fast で磨き、最高の精細さが必要なときは Standard で納品します。
- 03
アイデンティティを保つリファレンスから動画
google/veo3.1/reference-to-video は、参照画像から特定の被写体の見た目を保ちます。Fast 階層の版は、1回の生成につき最大3枚の参照画像に対応します。プレゼンター動画、ブランドキャラクター、シリーズもののコンテンツに便利です。
- 04
video-extend は7秒単位で動作する
google/veo3.1-fast/video-extend は、Veo 3.1 の動画を、一貫したモーションとスタイルを保って、7秒単位で延長します。出力は、後処理でつなぎ合わせる別々のセグメントではなく、1本に統合されたクリップです。
- 05
開始・終了の補間は Lite のみ
google/veo3.1-lite/start-end-to-video は、開始画像と任意の終了画像の間を補間します。720p / 1080p、横向きと縦向きに対応しています。Lite 専用で、Standard や Fast では利用できません。
- 06
音のない映像ではオーディオをオフにする
generate_audio パラメーターは、Standard のテキストから動画エンドポイントで公開されています。Bロール、SNS投稿、どうせ後処理でオーディオを差し替えるクリップでは、オフにしてください。
料金
Veo 3.1 APIの料金
料金は出力ごとです。最終的な請求額は、各バリアントのプレイグラウンドで設定するパラメーター(解像度、長さ、出力数、参照入力)に応じて変わります。
| エンドポイント | タイプ | 最低料金 |
|---|---|---|
| google/ | text-to-video | $0.30 |
| google/ | video-extend | $3.20 |
| google/ | reference-to-video | $0.64 |
| google/ | video-extend | $1.20 |
| google/ | text-to-video | $1.20 |
| google/ | reference-to-video | $3.20 |
| google/ | image-to-video | $0.40 |
| google/ | text-to-video | $3.20 |
| google/ | image-to-video | $0.30 |
| google/ | image-to-video | $1.20 |
| google/ | image-to-video | $3.20 |
API
Veo 3.1 APIを呼び出す
wavespeed.ai/accesskeyでAPIキーに登録し、RESTで予測を送信します。プレイグラウンドでは、入力の組み合わせに応じて、そのまま貼り付けられるサンプルが生成されます。
POSThttps://api.wavespeed.ai/api/v3/google/veo3.1/text-to-video
# 1. Submit the prediction.
SUBMIT_RESPONSE=$(curl --silent --show-error --fail-with-body \
-X POST "https://api.wavespeed.ai/api/v3/google/veo3.1/text-to-video" \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $WAVESPEED_API_KEY" \
-d '{
"prompt": "A cinematic shot of a city at sunset, soft golden light",
"aspect_ratio": "16:9",
"duration": 8,
"resolution": "1080p",
"generate_audio": true
}')
TASK=$(printf '%s' "$SUBMIT_RESPONSE" | jq 'if has("data") then .data else . end')
PREDICTION_ID=$(printf '%s' "$TASK" | jq -r '.id')
if [ -z "$PREDICTION_ID" ] || [ "$PREDICTION_ID" = "null" ]; then
printf 'Submission response did not contain a prediction id
' >&2
exit 1
fi
RESULT_URL="https://api.wavespeed.ai/api/v3/predictions/$PREDICTION_ID/result"
# 2. Poll until the prediction finishes.
while true; do
RESPONSE=$(curl --silent --show-error --fail-with-body "$RESULT_URL" \
-H "Authorization: Bearer $WAVESPEED_API_KEY")
RESULT=$(printf '%s' "$RESPONSE" | jq 'if has("data") then .data else . end')
STATUS=$(printf '%s' "$RESULT" | jq -r '.status')
case "$STATUS" in
completed) printf '%s\n' "$RESULT" | jq '.outputs'; break ;;
failed|cancelled|timeout|deleted) printf '%s\n' "$RESULT" | jq . >&2; exit 1 ;;
*) sleep 2 ;;
esac
doneconst submitUrl = "https://api.wavespeed.ai/api/v3/google/veo3.1/text-to-video";
const apiKey = process.env.WAVESPEED_API_KEY;
if (!apiKey) throw new Error('Set WAVESPEED_API_KEY');
async function requestJson(url, options = {}) {
const response = await fetch(url, options);
if (!response.ok) throw new Error(await response.text());
return response.json();
}
// 1. Submit the prediction.
const body = await requestJson(submitUrl, {
method: "POST",
headers: {
"Authorization": `Bearer ${apiKey}`,
"Content-Type": "application/json",
},
body: JSON.stringify({
"prompt": "A cinematic shot of a city at sunset, soft golden light",
"aspect_ratio": "16:9",
"duration": 8,
"resolution": "1080p",
"generate_audio": true
}),
});
const task = body.data ?? body;
const resultUrl = `https://api.wavespeed.ai/api/v3/predictions/${task.id}/result`;
// 2. Poll until the prediction finishes.
while (true) {
const resultBody = await requestJson(resultUrl, {
headers: { "Authorization": `Bearer ${apiKey}` },
});
const result = resultBody.data ?? resultBody;
if (result.status === "completed") {
console.log(result.outputs);
break;
}
if (["failed", "cancelled", "timeout", "deleted"].includes(result.status)) throw new Error(JSON.stringify(result));
await new Promise(resolve => setTimeout(resolve, 2000));
}import json
import os
import time
from urllib.request import Request, urlopen
api_key = os.environ["WAVESPEED_API_KEY"]
headers = {"Authorization": f"Bearer {api_key}", "Content-Type": "application/json"}
payload = {
"prompt": "A cinematic shot of a city at sunset, soft golden light",
"aspect_ratio": "16:9",
"duration": 8,
"resolution": "1080p",
"generate_audio": True
}
def request_json(url, data=None):
request = Request(url, data=data, headers=headers, method="POST" if data else "GET")
with urlopen(request) as response:
return json.load(response)
# 1. Submit the prediction.
body = request_json("https://api.wavespeed.ai/api/v3/google/veo3.1/text-to-video", json.dumps(payload).encode())
task = body.get("data", body)
result_url = f"https://api.wavespeed.ai/api/v3/predictions/{task['id']}/result"
# 2. Poll until the prediction finishes.
while True:
result_body = request_json(result_url)
result = result_body.get("data", result_body)
status = result.get("status")
if status == "completed":
print(result.get("outputs", []))
break
if status in {"failed", "cancelled", "timeout", "deleted"}:
raise RuntimeError(result)
time.sleep(2)比較
Veo 3.1と他モデルの比較
WaveSpeedAI上の類似モデルではなくVeo 3.1を選ぶべきケース。
Veo 3.1 vs Seedance 2.0
Seedance 2.0 は、全階層でのネイティブオーディオと高速な Turbo 階層を備えています。Veo 3.1 は最上位階層では高価ですが、Lite 階層はコスト面で競争力があり、人物の顔については Veo のフォトリアリズムの評価が高いです。
Veo 3.1 vs Kling 3.0
Kling 3.0 には Pro と 4K の階層とモーションコントロールエンドポイントがあります。Veo 3.1 は3つの料金階層(Standard / Fast / Lite)を備え、リファレンスから動画と開始・終了の補間を主要なエンドポイントとして提供しています。機能の構成が異なります。
Veo 3.1 vs Wan 2.7
Wan 2.7 は、リファレンスから動画、画像編集、テキストから画像のバリアントを1つのファミリーに備え、モダリティをまたぐツールが幅広くなっています。Veo 3.1 の Standard 階層は、Lite から Standard までの階層によるコスト最適化と、Wan とは扱いが異なる7秒の video-extend をカバーします。
FAQ
Veo 3.1 API — よくある質問
料金、ライセンス、連携など、WaveSpeedAIでVeo 3.1を実行する際のよくある質問。
Veo 3.1 APIとは何ですか?
Veo 3.1は、Googleの動画生成モデルで、WaveSpeedAI上でREST APIとして提供されています。Google の Veo 3.1 は、1080pの同期ネイティブオーディオ付きのテキストから動画に対応しています。3つの階層(Standard、Fast、Lite)に、テキストから動画、画像から動画、リファレンスから動画、video-extend があり、Lite 階層では start-end-to-video も利用できます。プログラムから呼び出すことも、上にリンクされたプレイグラウンドで試すこともできます。
Veo 3.1 APIはどう呼び出しますか?
WaveSpeedAIのアカウントに登録し、/accesskeyからAPIキーをコピーして、入力をJSONとしてhttps://api.wavespeed.ai/api/v3/google/veo3.1/text-to-videoにPOSTします。エンドポイントは予測IDを返します。結果のエンドポイントを約2秒ごとにポーリングし、長時間かかるタスクでは間隔を広げ、終端ステータスになったら停止してください。本番運用向けのPython / Node.js / cURLのサンプルは上にあります。
Veo 3.1 APIの料金はいくらですか?
Veo 3.1は1回あたり$0.30からです。実際の料金は、設定するパラメーター(解像度、長さ、出力数、参照入力)に応じて変わります。プレイグラウンドの「生成」ボタンの横に表示されるリアルタイムの料金プレビューで、現在の入力での正確な料金を確認できます。
Veo 3.1にはどのバリアントがありますか?
WaveSpeedAIでは、11個のVeo 3.1エンドポイントが利用可能です:google/veo3.1-lite/text-to-video, google/veo3.1/video-extend, google/veo3.1-fast/reference-to-video, google/veo3.1-fast/video-extend, google/veo3.1-fast/text-to-video, google/veo3.1/reference-to-video, google/veo3.1-lite/start-end-to-video, google/veo3.1/text-to-videoほか。各バリアントには、専用のプレイグラウンドページと料金があります。
Veo 3.1の出力を商用利用できますか?
商用利用の権利は、Googleのモデルライセンスに従います。Googleのほとんどのモデルは出力の商用利用を認めています。具体的なライセンスの概要は各モデルのプレイグラウンドページで、プラットフォーム全体の条件はWaveSpeedAIの利用規約でご確認ください。
なぜ直接ではなくWaveSpeedAIでVeo 3.1を使うのですか?
Veo 3.1と、他のプロバイダーの1,000以上のAIモデルを、1つのAPIキー、1つの請求アカウントで利用できます。ベンダーごとのSDK設定も、個別のレート制限も、ベンダーごとの連携コードの書き直しも不要です。料金は、通常Googleの直接のAPIと同等かそれ以下です。
提供元
Googleについて
WaveSpeedAIにおけるVeo 3.1と、Googleのモデルラインナップ全体を手がけるチーム。
Google のAI研究は、主に Google DeepMind と Google Research で行われています。画像・動画モデルである Imagen、Veo、そして Nano Banana(Gemini 3 Image)のような Gemini 系のマルチモーダルモデルは、より広い Gemini ラインナップとアーキテクチャや学習基盤を共有しています。出力は、正確な文字描画、幅広いスタイルへの対応、商用利用に適したライセンスで知られています。
WaveSpeedAIでVeo 3.1を使った開発を始めよう
登録時に無料のスタータークレジットを進呈。Googleをはじめ、あらゆるプロバイダーの1,000以上のAIモデルを、1つのAPIキーで。