Veo 3.1 API
Google Veo 3.1 — 1080p에서 동기화된 네이티브 오디오를 지원하는 텍스트-투-비디오입니다. Standard, Fast, Lite의 세 등급에서 텍스트-투-비디오, 이미지-투-비디오, 레퍼런스-투-비디오, video-extend를 제공하며, Lite 등급에는 start-end-to-video도 있습니다.
Standard 등급은 Veo의 최고 품질을, Fast 등급은 반복 작업을, Lite 등급은 프리비즈와 대량 작업을 위한 것입니다. 레퍼런스-투-비디오는 정체성을 유지하는 생성을 위해 레퍼런스 이미지를 받으며, video-extend는 7초 단위로 이어 하나의 병합된 클립을 만듭니다.
개요
Veo 3.1 API 소개
Veo 3.1의 기능, Google 모델 라인업에서의 위치, 그리고 팀들이 이 모델을 선택하는 이유.
Google의 비디오 생성 모델 Veo 3.1. WaveSpeedAI REST API로 바로 사용할 수 있습니다. Google Veo 3.1 — 1080p에서 동기화된 네이티브 오디오를 지원하는 텍스트-투-비디오입니다. Standard, Fast, Lite의 세 등급에서 텍스트-투-비디오, 이미지-투-비디오, 레퍼런스-투-비디오, video-extend를 제공하며, Lite 등급에는 start-end-to-video도 있습니다.
Standard 등급은 Veo의 최고 품질을, Fast 등급은 반복 작업을, Lite 등급은 프리비즈와 대량 작업을 위한 것입니다. 레퍼런스-투-비디오는 정체성을 유지하는 생성을 위해 레퍼런스 이미지를 받으며, video-extend는 7초 단위로 이어 하나의 병합된 클립을 만듭니다.
WaveSpeedAI의 Veo 3.1 제품군은 Text-To-Video, Video-Extend, Reference-To-Video, Image-To-Video 워크플로를 아우르는 REST 엔드포인트 11개를 제공합니다. 각 변형은 고유한 가격, 파라미터, 예시 출력을 갖고 있으니 입력 방식과 운영 조건에 맞는 것을 고르거나, 같은 API 키로 여러 개를 호출해 다단계 파이프라인을 구성하세요.
Veo 3.1 실행에도 WaveSpeedAI의 다른 1,000개 이상의 AI 모델과 같은 API 키, 결제 계정, 요청 한도 체계를 그대로 사용합니다. 별도의 벤더 설정도, 제공사별 SDK도, 벤더별 요청 한도 체계도 필요 없습니다. 하나의 연동으로 텍스트 → 이미지, 텍스트 → 비디오부터 오디오 합성, 3D 생성, 업스케일링, 편집까지 모두 지원합니다.
엔드포인트
Veo 3.1 API 전체 엔드포인트
Veo 3.1 엔드포인트 11개를 지금 WaveSpeedAI에서 사용할 수 있습니다. 워크플로에 맞는 변형을 고르세요.
/filters:quality(82)/media/images/20260408104718_8uxbq2i4.webp)
Veo3.1 Lite Text To Video
Google Veo 3.1 Lite Text-to-Video generates high-fidelity 720p or 1080p videos with natively generated audio from text prompts. Lightweight variant optimized for cost efficiency. Supports landscape and portrait aspect ratios, dialogue with lip-sync, and customizable duration. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
/filters:quality(82)/media/images/20260408104740_drl8tu9l.webp)
Veo3.1 Video Extend
Extend and continue Veo 3.1 videos with smooth motion, preserved style, and strong scene coherence. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
/filters:quality(82)/media/images/1778749312530573096_makufpyI.webp)
Veo3.1 Fast Reference To Video
Google Veo 3.1 Fast Reference to Video is a fast AI reference-to-video generation model that creates 8-second videos from up to three reference images using the official Veo predictLongRunning endpoint with referenceImages assets. Ready-to-use REST inference API for product videos, character consistency, branded visual storytelling, social media clips, advertising creatives, and professional reference-based video generation workflows with simple integration, no coldstarts, and affordable pricing.
/filters:quality(82)/media/images/20260408104757_ct9o5p9o.webp)
Veo3.1 Fast Video Extend
Extend Veo 3.1 videos in 7-second steps with the Fast endpoint—quick, coherent continuation that preserves style and motion, output as a single merged clip. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
/filters:quality(82)/media/images/20260408104805_tr00kyv7.webp)
Veo3.1 Fast Text To Video
Google Veo 3.1 Fast creates text-to-video with native 1080p and synchronized audio, delivering high-quality videos for creators. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
/filters:quality(82)/media/images/20260408104748_nwaixdpw.webp)
Veo3.1 Reference To Video
Google Veo3.1 Reference-to-Video performs image-to-video generation that preserves a specific subject's appearance and identity from provided reference images, enabling consistent character or product motion across frames. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
/filters:quality(82)/media/images/20260408104714_ut4h5kku.webp)
Veo3.1 Lite Start End To Video
Google Veo 3.1 Lite Start-End-to-Video generates high-fidelity videos by interpolating between a start image and an optional end image. Supports 720p and 1080p resolutions, landscape and portrait aspect ratios, and native audio generation. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
/filters:quality(82)/media/images/20260408104752_1voe0l4b.webp)
Veo3.1 Text To Video
Google Veo 3.1 converts text prompts into videos with synchronized audio at native 1080p for high-quality outputs. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
/filters:quality(82)/media/images/20260408104708_dmnbefn0.webp)
Veo3.1 Lite Image To Video
Google Veo 3.1 Lite Image-to-Video transforms static images into high-fidelity 720p or 1080p videos with natively generated audio. Supports many interpolation use cases, landscape and portrait aspect ratios, and customizable duration. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
/filters:quality(82)/media/images/20260408104802_aaez2mmf.webp)
Veo3.1 Fast Image To Video
Google Veo 3.1 Fast is an Image-to-Video model with native 1080p output for high-detail videos from images and fast performance. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
/filters:quality(82)/media/images/20260408104744_xfvl07s5.webp)
Veo3.1 Image To Video
Google Veo 3.1 is an Image-to-Video model that converts images into high-quality videos with native 1080P output for enhanced detail and creative flexibility. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
예시
Veo 3.1 활용 예시
Veo 3.1 API로 생성한 실제 출력물입니다. 비디오 위에 마우스를 올리면 미리 보고, 클릭하면 전체 크기 뷰어가 열립니다.
사용 방법
Veo 3.1 API 사용 방법
가입부터 결과물 생성까지 네 단계입니다. Python, Node.js, cURL 전체 예시는 아래 API 섹션에 있습니다.
- 01
API 키 받기
WaveSpeedAI 계정을 만들고 대시보드에서 API 키를 복사하세요. 신규 계정에는 무료 스타터 크레딧이 제공되어, 요금이 청구되기 전에 Playground를 수십 번 실행해 볼 수 있습니다.
- 02
Prediction 제출
입력을 JSON으로 https://api.wavespeed.ai/api/v3/google/veo3.1/text-to-video에 POST하세요. 엔드포인트가 prediction id를 즉시 반환합니다. 생성은 비동기로 처리되므로 추론 중에 연결을 계속 열어 둘 필요가 없습니다.
- 03
완료될 때까지 폴링
https://api.wavespeed.ai/api/v3/predictions/{request_id}/result로 GET 요청을 보내세요. completed이면 outputs를 반환하고, failed, cancelled, timeout, deleted이면 오류로 종료하며, 그 밖의 모든 status에서는 폴링을 계속합니다.
- 04
출력 URL 읽기
status가 "completed"가 되면 data.outputs[0]에서 URL을 읽으세요. 이 URL은 WaveSpeedAI CDN에 있는 생성된 미디어를 가리키며, 호출한 Veo 3.1 변형에 따라 이미지, 비디오, 오디오 또는 3D 파일입니다.
활용 사례
Veo 3.1 활용 사례
개발자와 크리에이터가 Veo 3.1 API를 사용하는 대표적인 워크플로입니다.
네이티브 1080p 오디오를 지원하는 텍스트-투-비디오
google/veo3.1/text-to-video는 텍스트 프롬프트를 네이티브 1080p에서 동기화된 오디오가 포함된 비디오로 변환합니다. 업스케일링 없이 바로 납품 가능한 해상도입니다. 같은 프롬프트 형식이 Standard, Fast, Lite 등급 모두에서 작동합니다.
정체성 유지를 위한 레퍼런스-투-비디오
google/veo3.1/reference-to-video는 제공된 레퍼런스 이미지에서 특정 피사체의 외형을 유지하면서 이미지-투-비디오를 수행합니다. Fast 등급 레퍼런스-투-비디오는 Veo의 공식 predictLongRunning 엔드포인트를 사용하며, 8초 출력에 레퍼런스 이미지를 최대 3장까지 지원합니다.
등급별 가격: Standard / Fast / Lite
납품에는 Standard, 반복 작업에는 Fast, 프리비즈와 대량 작업에는 Lite를 사용하세요. Lite는 품질 면에서 절충이 있으며, 프롬프트를 다시 쓰지 않고 엔드포인트 URL만 바꿔 등급을 전환할 수 있습니다.
7초 단위의 video-extend
google/veo3.1/video-extend는 기존 Veo 클립을 부드러운 모션과 유지된 스타일로 이어 줍니다. 결과물은 따로 이어 붙여야 하는 개별 구간이 아니라 하나로 병합된 클립입니다. Fast 등급 video-extend(/extension)는 반복적인 연장에 더 저렴한 옵션입니다.
시작-종료 보간(Lite 등급)
google/veo3.1-lite/start-end-to-video는 시작 이미지와 선택적인 종료 이미지 사이를 보간해 연결되는 모션을 생성합니다. 720p와 1080p, 가로 및 세로 화면 비율을 지원합니다. Lite 가격대에서 애니매틱과 키프레임 기반 워크플로에 유용합니다.
네이티브 1080p 이미지-투-비디오
google/veo3.1/image-to-video는 이미지를 향상된 디테일과 창의적 유연성을 갖춘 1080p 비디오로 변환합니다. 텍스트만의 브리프가 아니라 핵심 스틸에서 시작할 때 알맞은 선택입니다.
팁
Veo 3.1 프롬프트 작성 팁
Veo 3.1에서 더 나은 결과를 얻기 위한 실용적인 조언입니다. 운영 파이프라인에서 비디오 모델 전반에 통하는 패턴을 바탕으로 했습니다.
- 01
모든 등급에서 네이티브 1080p 출력
카탈로그 설명: "고품질 출력을 위한 네이티브 1080p". 모든 Veo 3.1 변형은 1080p로 바로 생성하므로 납품 전에 업스케일링 패스가 필요 없습니다. 대역폭이 더 중요할 때는 Lite 등급에서 720p도 지원합니다.
- 02
Lite로 반복하고, Fast로 다듬고, Standard로 납품하세요
같은 프롬프트 형식이 Standard, Fast, Lite 모두에서 작동하므로 다시 쓰지 않고 엔드포인트 URL만 바꾸면 됩니다. 프롬프트 방향은 Lite로 반복하고, 더 높은 품질은 Fast로 다듬고, 최고 충실도가 중요할 때는 Standard로 납품하세요.
- 03
정체성 유지를 위한 레퍼런스-투-비디오
google/veo3.1/reference-to-video는 레퍼런스 이미지에서 특정 피사체의 외형을 유지합니다. Fast 등급 버전은 생성당 최대 3장의 레퍼런스 이미지를 지원합니다. 발표자 영상, 브랜드 캐릭터, 연재 콘텐츠에 유용합니다.
- 04
video-extend는 7초 단위로 작동합니다
google/veo3.1-fast/video-extend는 Veo 3.1 비디오를 일관된 모션과 유지된 스타일로 7초 단위로 연장합니다. 결과물은 후반 작업에서 이어 붙여야 하는 개별 구간이 아니라 하나로 병합된 클립입니다.
- 05
시작-종료 보간은 Lite에서만 가능합니다
google/veo3.1-lite/start-end-to-video는 시작 이미지와 선택적인 종료 이미지 사이를 보간합니다. 720p / 1080p, 가로와 세로를 지원합니다. Lite 전용이며 Standard나 Fast에서는 사용할 수 없습니다.
- 06
무음 영상에는 오디오를 끄세요
generate_audio 파라미터는 Standard 텍스트-투-비디오 엔드포인트에서 제공됩니다. B-roll, 소셜 게시물, 어차피 후반 작업에서 오디오를 교체할 클립에서는 끄세요.
가격
Veo 3.1 API 가격
가격은 출력당 책정됩니다. 최종 요금은 각 변형의 Playground에서 설정한 파라미터(해상도, 길이, 출력 수, 레퍼런스)에 따라 달라집니다.
| 엔드포인트 | 유형 | 시작 가격 |
|---|---|---|
| google/ | text-to-video | $0.30 |
| google/ | video-extend | $3.20 |
| google/ | reference-to-video | $0.64 |
| google/ | video-extend | $1.20 |
| google/ | text-to-video | $1.20 |
| google/ | reference-to-video | $3.20 |
| google/ | image-to-video | $0.40 |
| google/ | text-to-video | $3.20 |
| google/ | image-to-video | $0.30 |
| google/ | image-to-video | $1.20 |
| google/ | image-to-video | $3.20 |
API
Veo 3.1 API 호출하기
wavespeed.ai/accesskey에서 API 키를 발급받은 뒤 REST로 prediction을 제출하세요. Playground에서 입력 조합별로 바로 붙여넣을 수 있는 샘플이 생성됩니다.
POSThttps://api.wavespeed.ai/api/v3/google/veo3.1/text-to-video
# 1. Submit the prediction.
SUBMIT_RESPONSE=$(curl --silent --show-error --fail-with-body \
-X POST "https://api.wavespeed.ai/api/v3/google/veo3.1/text-to-video" \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $WAVESPEED_API_KEY" \
-d '{
"prompt": "A cinematic shot of a city at sunset, soft golden light",
"aspect_ratio": "16:9",
"duration": 8,
"resolution": "1080p",
"generate_audio": true
}')
TASK=$(printf '%s' "$SUBMIT_RESPONSE" | jq 'if has("data") then .data else . end')
PREDICTION_ID=$(printf '%s' "$TASK" | jq -r '.id')
if [ -z "$PREDICTION_ID" ] || [ "$PREDICTION_ID" = "null" ]; then
printf 'Submission response did not contain a prediction id
' >&2
exit 1
fi
RESULT_URL="https://api.wavespeed.ai/api/v3/predictions/$PREDICTION_ID/result"
# 2. Poll until the prediction finishes.
while true; do
RESPONSE=$(curl --silent --show-error --fail-with-body "$RESULT_URL" \
-H "Authorization: Bearer $WAVESPEED_API_KEY")
RESULT=$(printf '%s' "$RESPONSE" | jq 'if has("data") then .data else . end')
STATUS=$(printf '%s' "$RESULT" | jq -r '.status')
case "$STATUS" in
completed) printf '%s\n' "$RESULT" | jq '.outputs'; break ;;
failed|cancelled|timeout|deleted) printf '%s\n' "$RESULT" | jq . >&2; exit 1 ;;
*) sleep 2 ;;
esac
doneconst submitUrl = "https://api.wavespeed.ai/api/v3/google/veo3.1/text-to-video";
const apiKey = process.env.WAVESPEED_API_KEY;
if (!apiKey) throw new Error('Set WAVESPEED_API_KEY');
async function requestJson(url, options = {}) {
const response = await fetch(url, options);
if (!response.ok) throw new Error(await response.text());
return response.json();
}
// 1. Submit the prediction.
const body = await requestJson(submitUrl, {
method: "POST",
headers: {
"Authorization": `Bearer ${apiKey}`,
"Content-Type": "application/json",
},
body: JSON.stringify({
"prompt": "A cinematic shot of a city at sunset, soft golden light",
"aspect_ratio": "16:9",
"duration": 8,
"resolution": "1080p",
"generate_audio": true
}),
});
const task = body.data ?? body;
const resultUrl = `https://api.wavespeed.ai/api/v3/predictions/${task.id}/result`;
// 2. Poll until the prediction finishes.
while (true) {
const resultBody = await requestJson(resultUrl, {
headers: { "Authorization": `Bearer ${apiKey}` },
});
const result = resultBody.data ?? resultBody;
if (result.status === "completed") {
console.log(result.outputs);
break;
}
if (["failed", "cancelled", "timeout", "deleted"].includes(result.status)) throw new Error(JSON.stringify(result));
await new Promise(resolve => setTimeout(resolve, 2000));
}import json
import os
import time
from urllib.request import Request, urlopen
api_key = os.environ["WAVESPEED_API_KEY"]
headers = {"Authorization": f"Bearer {api_key}", "Content-Type": "application/json"}
payload = {
"prompt": "A cinematic shot of a city at sunset, soft golden light",
"aspect_ratio": "16:9",
"duration": 8,
"resolution": "1080p",
"generate_audio": True
}
def request_json(url, data=None):
request = Request(url, data=data, headers=headers, method="POST" if data else "GET")
with urlopen(request) as response:
return json.load(response)
# 1. Submit the prediction.
body = request_json("https://api.wavespeed.ai/api/v3/google/veo3.1/text-to-video", json.dumps(payload).encode())
task = body.get("data", body)
result_url = f"https://api.wavespeed.ai/api/v3/predictions/{task['id']}/result"
# 2. Poll until the prediction finishes.
while True:
result_body = request_json(result_url)
result = result_body.get("data", result_body)
status = result.get("status")
if status == "completed":
print(result.get("outputs", []))
break
if status in {"failed", "cancelled", "timeout", "deleted"}:
raise RuntimeError(result)
time.sleep(2)비교
Veo 3.1 vs 대안
비슷한 WaveSpeedAI 모델과 비교한 Veo 3.1의 선택 기준.
Veo 3.1 vs Seedance 2.0
Seedance 2.0은 모든 등급에서 네이티브 오디오를 제공하고 빠른 Turbo 등급이 있습니다. Veo 3.1은 최상위 등급에서 더 비싸지만 Lite 등급은 비용 면에서 경쟁력이 있으며, 사람 얼굴 표현에서는 Veo의 사실적인 표현이라는 평판이 더 강합니다.
Veo 3.1 vs Kling 3.0
Kling 3.0은 Pro 및 4K 등급과 모션 컨트롤 엔드포인트를 갖추고 있습니다. Veo 3.1은 세 가지 가격 등급(Standard / Fast / Lite)을 제공하며 레퍼런스-투-비디오와 시작-종료 보간을 핵심 엔드포인트로 제공합니다. 기능 구성이 서로 다릅니다.
Veo 3.1 vs Wan 2.7
Wan 2.7은 한 패밀리에 레퍼런스-투-비디오, 이미지 편집, 텍스트-투-이미지 변형이 있어 교차 모달 도구 모음이 더 넓습니다. Veo 3.1의 Standard 등급은 Lite부터 Standard까지의 등급별 비용 최적화와, Wan이 다르게 처리하는 7초 단위 video-extend를 제공합니다.
FAQ
Veo 3.1 API — 자주 묻는 질문
가격, 라이선스, 연동 — WaveSpeedAI에서 Veo 3.1 실행에 관한 자주 묻는 질문.
Veo 3.1 API란 무엇인가요?
Veo 3.1 — WaveSpeedAI에서 REST API로 제공되는 Google의 비디오 생성 모델입니다. Google Veo 3.1 — 1080p에서 동기화된 네이티브 오디오를 지원하는 텍스트-투-비디오입니다. Standard, Fast, Lite의 세 등급에서 텍스트-투-비디오, 이미지-투-비디오, 레퍼런스-투-비디오, video-extend를 제공하며, Lite 등급에는 start-end-to-video도 있습니다. 프로그래밍 방식으로 호출하거나 위에 링크된 Playground에서 체험해 볼 수 있습니다.
Veo 3.1 API는 어떻게 호출하나요?
WaveSpeedAI 계정을 만들고 /accesskey에서 API 키를 복사한 뒤, 입력을 JSON으로 https://api.wavespeed.ai/api/v3/google/veo3.1/text-to-video에 POST하세요. 엔드포인트가 prediction id를 반환합니다. 결과 엔드포인트 폴링은 약 2초마다 시작하고, 오래 걸리는 작업은 간격을 늘리며, 종료 status가 나오면 중단하세요. 운영 환경용 Python / Node.js / cURL 예시는 위에 있습니다.
Veo 3.1 API 요금은 얼마인가요?
Veo 3.1의 가격은 회당 $0.30부터입니다. 정확한 비용은 설정한 파라미터(해상도, 길이, 출력 수, 레퍼런스)에 따라 달라집니다. Playground의 생성 버튼 옆 실시간 비용 미리보기에서 현재 입력 기준의 정확한 가격을 확인할 수 있습니다.
Veo 3.1에는 어떤 변형이 있나요?
WaveSpeedAI에는 라이브 Veo 3.1 엔드포인트가 11개 있습니다: google/veo3.1-lite/text-to-video, google/veo3.1/video-extend, google/veo3.1-fast/reference-to-video, google/veo3.1-fast/video-extend, google/veo3.1-fast/text-to-video, google/veo3.1/reference-to-video, google/veo3.1-lite/start-end-to-video, google/veo3.1/text-to-video 외 다수. 각 변형에는 고유한 Playground 페이지와 가격이 있습니다.
Veo 3.1 출력물을 상업적으로 사용할 수 있나요?
상업적 사용 권한은 Google 모델 라이선스를 따릅니다. 대부분의 Google 모델은 출력물의 상업적 사용을 허용합니다. 구체적인 라이선스 요약은 각 모델의 Playground 페이지를, 플랫폼 수준의 조건은 WaveSpeedAI 이용약관을 참고하세요.
직접 연동하지 않고 WaveSpeedAI에서 Veo 3.1 모델을 사용하는 이유는 무엇인가요?
API 키 하나와 결제 계정 하나로 Veo 3.1 모델과 다른 제공사의 1,000개 이상의 AI 모델을 모두 사용합니다. 벤더별 SDK 설정도, 별도의 요청 한도 체계도, 벤더마다 연동 코드를 다시 쓸 일도 없습니다. 가격은 대체로 Google의 직접 API와 같거나 더 저렴합니다.
제공사
Google 소개
WaveSpeedAI의 Veo 3.1 및 Google 전체 모델 라인업을 만든 팀.
Google의 AI 연구는 주로 Google DeepMind와 Google Research에서 이루어집니다. Imagen, Veo 같은 이미지 및 비디오 모델과 Nano Banana(Gemini 3 Image) 같은 Gemini 계열 멀티모달 모델은 더 넓은 Gemini 라인업과 아키텍처 및 학습 인프라를 공유합니다. 정확한 텍스트 렌더링, 폭넓은 스타일 범위, 상업용 수준의 라이선스로 평가받습니다.
WaveSpeedAI에서 Veo 3.1 시작하기
가입 시 무료 스타터 크레딧 제공. Google 및 다른 모든 제공사의 1,000개 이상 AI 모델을 API 키 하나로.