Bytedance Seedance 2.0 Video Edit API Documentation

Bytedance Seedance 2.0 Video Edit API Documentation

Playground

Try it on WaveSpeedAI!

Seedance 2.0 (Video-Edit) edits an input video from a natural-language prompt. The reference video drives subject identity, composition, and motion while the model rewrites lighting, style, weather, environment, or specific elements as instructed. Built on ByteDance Seed’s unified multimodal architecture for cinematic, motion-stable output. Ready-to-use REST API, best performance, no coldstarts, affordable pricing.

Features

Seedance 2.0 Video-Edit transforms an input video from a natural-language prompt — change lighting, weather, style, environment, or specific elements while preserving the subject identity, composition, and motion of the original. Built on ByteDance Seed’s unified multimodal architecture for cinematic, motion-stable output.


Key Features

  • Conversational video editing — Describe the change in plain language; the model rewrites the scene while keeping the original motion intact.
  • Subject and motion preservation — Faces, objects, and camera movement from the input video stay consistent through the edit.
  • Multi-reference support — Optionally guide style, character identity, or audio with reference images and audio clips.
  • Native audio synchronization — Generates synchronized audio in a single pass.
  • Cinematic output quality — Director-level lighting, framing, and motion stability inherited from Seedance 2.0.

Parameters

ParameterRequiredDescription
promptYesDescribe the edit.
videoYesInput video URL. Videos longer than 15 s are trimmed to 15 s.
reference_imagesNoOptional reference images for style or character guidance.
reference_audiosNoOptional reference audio for audio guidance.
durationNoOutput length in seconds (4-15). Auto-detected from the input video if not specified.
aspect_ratioNo16:9, 9:16, 4:3, 3:4, 1:1, 21:9. Adapts to the input if not specified.
resolutionNo480p, 720p (default), 1080p, or 4k.
generate_audioNoGenerate synchronized audio for the output video (default: true)
enable_web_searchNoEnable web search for real-time context.

How to Use

  1. Upload the input video. Anything longer than 15 s is trimmed to the first 15 s automatically.
  2. Write the edit prompt. Describe the change you want — the prefix Edit the input video. is added for you.
  3. (Optional) Add references. Reference images can constrain style or identity; reference audio can constrain the soundtrack.
  4. (Optional) Set a duration. Auto-detected from the input video length if not provided.
  5. Run. Receive the edited video with synchronized audio.

Writing Effective Prompts

Editing keeps your source video and changes only what you name. Be precise about the scope, and describe the change as from A to B.

State the change plainly

Name exactly what to add, remove, or modify — and use an editing trigger word (add, remove, replace, change to, modify) so the intent is unambiguous.

  • “Replace the man’s coffee cup with a book, keep everything else unchanged.”
  • “Remove the subtitles.”
  • “Change the background from a city street to a snowy forest.”

Use timestamps for partial edits

Confine a change to a time range so the rest of the clip is untouched.

  • “From 4-6s, change the man’s action from drinking coffee to waving; leave the rest unchanged.”

Reference images for edits

Attach images to specify a replacement, and bind each explicitly (see below).

  • “Replace the woman on the right with the person in @image1.”

Referencing uploaded assets

When you attach images, videos, or audio, bind each one explicitly in the prompt by its upload order — @image1, @video1, @audio1 — and say what it’s for. Don’t rely on labels drawn inside the image itself.

  • “The knight in @image1 walks through the castle in @image2.”
  • “Refer to @video1 for the camera movement only; keep its shot order.”
  • “@image1 and @image2 are Character 1, voiced by @audio1.”

When a reference is already accurate, just point to it — no need to re-describe it in detail.

Audio edits

You can add, remove, or modify audio too.

  • “Translate the dialogue to Spanish, keep lip movements matched, no subtitles.”
  • “Remove the background music; keep only ambient and action sounds.”

Negative control

Positive descriptions work best, but you can suppress subtitles and audio:

  • “No subtitles.”
  • “No BGM — ambient and action sounds only.”
  • “No audio.”

Weak vs. strong

prompt
weakchange the video
strongReplace the red car in the video with a black motorcycle, matching its motion and speed. Keep the road, lighting, and camera movement unchanged. From 2s onward, add faint dust kicked up behind the rear wheel. Keep the original engine audio.

Pricing

Billed per second across input duration + output duration. Input duration is clamped to the 2-15 s range (sources shorter than 2 s are padded with the last frame).

ResolutionPer second
480p$0.075
720p$0.15
1080p$0.375
4k$0.750

Examples (input 5 s, output 5 s = 10 billed seconds):

ResolutionCost
480p$0.75
720p$1.50
1080p$3.75
4k$7.50

Examples (input 12 s, output 12 s = 24 billed seconds):

ResolutionCost
480p$1.80
720p$3.60
1080p$9.00
4k$18.00

Best Use Cases

  • Style and look transfer — Re-grade footage into a cinematic, vintage, animated, or stylized look.
  • Lighting and weather edits — Change time of day, add rain or snow, swap golden hour for blue hour.
  • Object or background swaps — Replace clothing, props, or environments while keeping motion intact.
  • Marketing variants — Generate quick variations of an existing ad clip without reshooting.

Pro Tips

  • Be specific about what should change and what should stay the same.
  • Mention lighting, mood, color palette, and camera intent for stronger output.
  • Use reference images when you need to lock in a particular character or style.
  • Trim your source video to the most relevant 4-15 s before uploading for the strongest edits.

Notes

  • Inputs longer than 15 s are trimmed to 15 s and inputs shorter than 2 s are padded with the last frame to 2 s before editing — billing reflects the conformed length.
  • Auto-detected duration matches the input length (rounded up), clamped to the 4-15 s range.
  • Native audio generation is included.


Notice: Use @image1, @image2, @audio1, etc. to reference your uploaded assets. The references will stay as plain text—don’t worry.


Authentication

For authentication details, please refer to the Authentication Guide.

API Endpoints

Submit Task & Query Result

set -euo pipefail

export WAVESPEED_API_KEY="your-api-key"

REQUEST_BODY=$(cat <<'JSON'
{
  "prompt": "A cinematic ocean wave at sunrise, highly detailed",
  "video": "https://interactive-examples.mdn.mozilla.net/media/cc0-videos/flower.mp4",
  "aspect_ratio": "16:9",
  "resolution": "720p",
  "enable_web_search": false,
  "generate_audio": true
}
JSON
)

# 1. Submit the prediction.
SUBMIT_RESPONSE=$(curl --silent --show-error --fail-with-body \
  -X POST "https://api.wavespeed.ai/api/v3/bytedance/seedance-2.0/video-edit" \
  -H "Authorization: Bearer ${WAVESPEED_API_KEY}" \
  -H "Content-Type: application/json" \
  -d "${REQUEST_BODY}")

TASK=$(printf '%s' "${SUBMIT_RESPONSE}" | jq 'if type == "object" and has("data") then .data else . end')
PREDICTION_ID=$(printf '%s' "${TASK}" | jq -r '.id // empty')
if [ -z "${PREDICTION_ID}" ]; then
  printf 'Submission response did not contain a prediction id
' >&2
  exit 1
fi
RESULT_URL="https://api.wavespeed.ai/api/v3/predictions/${PREDICTION_ID}/result"

# 2. Poll until the prediction finishes.
while true; do
  RESPONSE=$(curl --silent --show-error --fail-with-body \
    "${RESULT_URL}" \
    -H "Authorization: Bearer ${WAVESPEED_API_KEY}")
  RESULT=$(printf '%s' "${RESPONSE}" | jq 'if type == "object" and has("data") then .data else . end')
  STATUS=$(printf '%s' "${RESULT}" | jq -r '.status // empty')

  case "${STATUS}" in
    completed) printf '%s\n' "${RESULT}" | jq '.outputs'; break ;;
    failed|cancelled|timeout|deleted) printf '%s\n' "${RESULT}" | jq . >&2; exit 1 ;;
    *) sleep 2 ;;
  esac
done

Parameters

Task Submission Parameters

Request Parameters

ParameterTypeRequiredDefaultRangeDescription
promptstringYes-Describe the edit you want applied to the input video.
videostringYes-URL of the input video to edit. Videos longer than 15s are trimmed to 15s.
reference_imagesarray<string>No-0 ~ 9 itemsOptional reference image URLs to guide the edit (subject identity, style, etc.).
reference_audiosarray<string>No-0 ~ 3 itemsOptional reference audio URLs to guide audio generation.
aspect_ratiostringNo-16:9, 9:16, 4:3, 3:4, 1:1, 21:9Aspect ratio of the output video. Adapts to the input if not specified.
resolutionstringNo720p480p, 720p, 1080p, 4kOutput video resolution.
enable_web_searchbooleanNofalse-Enable web search for real-time information.
generate_audiobooleanNotrue-Whether to generate native audio for the edited output. Defaults to true. When set to false, the input video's audio track is preserved on the output instead.

Response Parameters

ParameterTypeDescription
codeintegerHTTP status code (e.g., 200 for success)
messagestringStatus message (e.g., “success”)
data.idstringUnique identifier for the prediction, Task Id
data.modelstringModel ID used for the prediction
data.outputsarrayOutput values, usually URL strings; some models return text strings or structured result objects (empty when status is not completed)
data.urlsobjectObject containing related API endpoints
data.statusstringTask status. completed is successful; failed, cancelled, timeout, and deleted are failure terminal statuses.
data.created_atstringISO timestamp of when the request was created (e.g., “2023-04-01T12:34:56.789Z”)
data.errorstringError message (empty if no error occurred)
data.timingsobjectObject containing timing details
data.timings.inferenceintegerInference time in milliseconds

Result Request Parameters

ParameterTypeRequiredDefaultDescription
idstringYes-Task ID

Result Response Parameters

ParameterTypeDescription
codeintegerHTTP status code (e.g., 200 for success)
messagestringStatus message (e.g., “success”)
dataobjectThe prediction data object containing all details
data.idstringUnique identifier for the prediction
data.modelstringModel ID used for the prediction
data.outputsarray<string | object>Array of generated outputs (empty when status is not completed). Items are usually URL strings, but may be text strings or structured result objects, depending on the model.
data.urlsobjectObject containing related API endpoints
data.statusstringStatus: completed is successful; failed, cancelled, timeout, and deleted are failure terminal statuses
data.created_atstringISO timestamp of when the request was created
data.errorstringError message (empty if no error occurred)
data.timingsobjectObject containing timing details
data.timings.inferenceintegerInference time in milliseconds
© 2026 WaveSpeedAI. All rights reserved.