Skywork AI Skyreels V3 Pro Single Avatar API Documentation

Skywork AI Skyreels V3 Pro Single Avatar API Documentation

Playground

Try it on WaveSpeedAI!

SkyReels V3 Pro Single Avatar is a high-quality AI talking avatar video generation model that creates audio-driven avatar videos from one image, one audio file, and a motion prompt. Ready-to-use REST inference API for digital humans, virtual presenters, product explainers, marketing videos, education content, social media clips, and professional avatar video workflows with simple integration, no coldstarts, and affordable pricing.

Features

Skywork AI SkyReels V3 Pro Single Avatar generates a talking avatar video from a single reference image and an audio clip. It is designed for higher-quality avatar performance, with stronger realism, smoother facial animation, and more polished lip-sync than the Standard variant, making it suitable for digital presenters, spokesperson videos, short-form content, and character-driven speaking clips.


Why Choose This?

  • Higher-quality avatar generation The Pro variant is built for stronger visual quality, more natural expression, and more refined speaking performance.

  • Single-image avatar workflow Turn one portrait image into a speaking avatar video.

  • Audio-driven lip-sync Use an uploaded audio clip to drive speech timing and mouth movement.

  • Prompt-guided behavior Add a short prompt to influence expression, delivery style, or overall presentation.

  • Simple production workflow Upload an image, upload audio, write a prompt, and generate a polished talking avatar clip.

  • Production-ready API Suitable for virtual presenters, social content, marketing spokespeople, and short-form avatar media.


Parameters

ParameterRequiredDescription
promptYesText instruction describing the desired avatar behavior, style, or delivery.
imageYesReference image used as the avatar source.
audioYesAudio track used to drive the avatar’s speaking performance.

How to Use

  1. Upload your image — provide a clear portrait image of the person you want to animate.
  2. Upload your audio — use a clean audio clip to drive the speaking performance.
  3. Write your prompt — describe the desired motion, expression, or delivery style.
  4. Submit — run the model and download the generated avatar video.

Example Prompt

Let the man speak naturally with subtle head movement, calm expression, and realistic lip-sync.


Pricing

Pricing is based on the uploaded audio duration.

Audio DurationCost
5s$0.40
10s$0.80
15s$1.20

Billing Rules

  • Pricing is $0.08 per second
  • Total price = $0.08 × audio duration
  • Longer audio increases cost linearly

Best Use Cases

  • Talking portrait videos — Animate a single portrait into a speaking clip.
  • Digital spokesperson content — Create presenter-style videos for announcements, marketing, or onboarding.
  • Virtual presenters — Generate clean speaking-avatar videos for explainers and business content.
  • Short-form social media clips — Turn portraits and voice clips into speaking content quickly.
  • Narration-based character videos — Pair a portrait with speech audio for expressive delivery.

Pro Tips

  • Use a clear, front-facing portrait for better avatar stability and facial animation.
  • Upload clean audio for stronger lip-sync and more natural speaking performance.
  • Keep the prompt simple and focused on expression or delivery style.
  • Shorter audio clips are easier to test when iterating on quality.
  • Use the Pro variant when you want better realism and polish than the Standard workflow.

Notes

  • prompt, image, and audio are required.
  • Pricing depends on the uploaded audio duration.
  • A clean portrait and high-quality audio generally improve results.
  • This workflow is intended for single-avatar speaking video generation.

Duration limit

The maximum supported audio duration is 200 seconds. Longer audio is automatically trimmed to 200 seconds before processing.

Authentication

For authentication details, please refer to the Authentication Guide.

API Endpoints

Submit Task & Query Result

set -euo pipefail

export WAVESPEED_API_KEY="your-api-key"

REQUEST_BODY=$(cat <<'JSON'
{
  "prompt": "A cinematic ocean wave at sunrise, highly detailed",
  "image": "https://interactive-examples.mdn.mozilla.net/media/cc0-images/painted-hand-298-332.jpg",
  "audio": "https://interactive-examples.mdn.mozilla.net/media/cc0-audio/t-rex-roar.mp3"
}
JSON
)

# 1. Submit the prediction.
SUBMIT_RESPONSE=$(curl --silent --show-error --fail-with-body \
  -X POST "https://api.wavespeed.ai/api/v3/skywork-ai/skyreels-v3-pro/single-avatar" \
  -H "Authorization: Bearer ${WAVESPEED_API_KEY}" \
  -H "Content-Type: application/json" \
  -d "${REQUEST_BODY}")

TASK=$(printf '%s' "${SUBMIT_RESPONSE}" | jq 'if type == "object" and has("data") then .data else . end')
PREDICTION_ID=$(printf '%s' "${TASK}" | jq -r '.id // empty')
if [ -z "${PREDICTION_ID}" ]; then
  printf 'Submission response did not contain a prediction id
' >&2
  exit 1
fi
RESULT_URL=$(printf '%s' "${TASK}" | jq -r '.urls.get // empty')
if [ -z "${RESULT_URL}" ]; then RESULT_URL="https://api.wavespeed.ai/api/v3/predictions/${PREDICTION_ID}/result"; fi

# 2. Poll until the prediction finishes.
while true; do
  RESPONSE=$(curl --silent --show-error --fail-with-body \
    "${RESULT_URL}" \
    -H "Authorization: Bearer ${WAVESPEED_API_KEY}")
  RESULT=$(printf '%s' "${RESPONSE}" | jq 'if type == "object" and has("data") then .data else . end')
  STATUS=$(printf '%s' "${RESULT}" | jq -r '.status // empty')

  case "${STATUS}" in
    completed) printf '%s\n' "${RESULT}" | jq '.outputs'; break ;;
    failed|cancelled|timeout) printf '%s\n' "${RESULT}" | jq . >&2; exit 1 ;;
    created|processing) sleep 2 ;;
    *) printf 'Unexpected status: %s
' "${STATUS}" >&2; exit 1 ;;
  esac
done

Parameters

Task Submission Parameters

Request Parameters

ParameterTypeRequiredDefaultRangeDescription
promptstringYes-Text prompt describing the scene, action, camera, or avatar behavior.
imagestringYes-Input image URL.
audiostringYes--Input audio URL.

Response Parameters

ParameterTypeDescription
codeintegerHTTP status code (e.g., 200 for success)
messagestringStatus message (e.g., “success”)
data.idstringUnique identifier for the prediction, Task Id
data.modelstringModel ID used for the prediction
data.outputsarrayOutput values, usually URL strings; some models return text strings or structured result objects (empty when status is not completed)
data.urlsobjectObject containing related API endpoints
data.urls.getstringURL to retrieve the prediction result
data.statusstringStatus of the task: created, processing, completed, or failed
data.created_atstringISO timestamp of when the request was created (e.g., “2023-04-01T12:34:56.789Z”)
data.errorstringError message (empty if no error occurred)
data.timingsobjectObject containing timing details
data.timings.inferenceintegerInference time in milliseconds

Result Request Parameters

ParameterTypeRequiredDefaultDescription
idstringYes-Task ID

Result Response Parameters

ParameterTypeDescription
codeintegerHTTP status code (e.g., 200 for success)
messagestringStatus message (e.g., “success”)
dataobjectThe prediction data object containing all details
data.idstringUnique identifier for the prediction
data.modelstringModel ID used for the prediction
data.outputsarray<string | object>Array of generated outputs (empty when status is not completed). Items are usually URL strings, but may be text strings or structured result objects, depending on the model.
data.urlsobjectObject containing related API endpoints
data.urls.getstringURL to poll for the prediction result
data.statusstringStatus: created, processing, completed, or failed
data.created_atstringISO timestamp of when the request was created
data.errorstringError message (empty if no error occurred)
data.timingsobjectObject containing timing details
data.timings.inferenceintegerInference time in milliseconds
© 2026 WaveSpeedAI. All rights reserved.