SAM 3 Video RLE is a unified foundation model for prompt-based segmentation in video. Track and segment objects across frames using text, points, or boxes, returning RLE encoded masks for efficient processing. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
Idle
{}$0.05per run·~20 / $1
the man
the product
the man
the woman
SAM3 Video Segmentation RLE segments and tracks target objects across video frames, then returns both a preview video URL and RLE mask data in JSON format. Use text prompts, point prompts, box prompts, or a combination of them to generate frame-level masks for programmatic video workflows.
Video object segmentation
Segment people, objects, animals, clothing, vehicles, or other visible targets across video frames.
Frame-level object tracking
Track selected subjects consistently through the video.
RLE mask output
Receive compact Run-Length Encoded mask data for downstream processing.
Multiple prompt types
Use text prompts, point prompts, box prompts, or combined prompt guidance for more accurate targeting.
Multi-object targeting
Track multiple objects by listing them in the prompt, such as person, car, dog.
Optional mask visualization
Use apply_mask to generate a visual preview of the segmentation result.
| Parameter | Required | Description |
|---|---|---|
| video | Yes | Source video to segment. Upload a video or provide a public video URL. |
| prompt | Yes | Text description of the object or objects to segment. Use short labels such as person, car, or person, cloth. |
| point_prompts | No | Point coordinates used to identify the target object or region. |
| box_prompts | No | Bounding box coordinates used to identify the target object or region. |
| apply_mask | No | Whether to apply the mask visualization to the video output. |
the man, red car, or person, cloth.apply_mask when you want a preview video with the segmentation mask applied.The response returns both:
RLE means Run-Length Encoding. Instead of storing every mask pixel one by one, it stores continuous runs of pixels, making mask data smaller and easier to pass into downstream computer-vision or editing pipelines.
The RLE JSON can be decoded with standard mask-processing tools when you need binary masks for each frame.
Pricing is based on input video duration.
Billed duration is rounded up to the next whole second, with a minimum billed duration of 3 seconds and a maximum billed duration of 600 seconds. Pricing is charged at $0.05 per 5 seconds.
| Billed Duration | Cost |
|---|---|
| 3s | $0.03 |
| 5s | $0.05 |
| 10s | $0.10 |
| 60s | $0.60 |
| 300s | $3.00 |
| 600s | $6.00 |
prompt, point_prompts, box_prompts, and apply_mask do not add separate charges.
person, dog, car, or shirt.person, backpack, bicycle.apply_mask when you want a visual preview of the segmentation result.video and prompt are required.Grab a WaveSpeedAI API key, then call POST https://api.wavespeed.ai/api/v3/wavespeed-ai/sam3-video-rle with your input as JSON. The endpoint returns a prediction id. Start polling the result endpoint around every 2 seconds, increase the interval for long-running tasks, and stop on any terminal status. On completed, read output values from data.outputs. Examples for Sam3 Video Rle below.
set -euo pipefail
: "${WAVESPEED_API_KEY:?Set WAVESPEED_API_KEY}"
REQUEST_BODY=$(cat <<'JSON'
{
"prompt": "A cinematic shot of a city at sunset, soft golden light",
"video": "https://interactive-examples.mdn.mozilla.net/media/cc0-videos/flower.mp4",
"apply_mask": true
}
JSON
)
# 1. Submit the prediction.
SUBMIT_RESPONSE=$(curl --silent --show-error --fail-with-body \
-X POST "https://api.wavespeed.ai/api/v3/wavespeed-ai/sam3-video-rle" \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $WAVESPEED_API_KEY" \
-d "$REQUEST_BODY")
TASK=$(printf '%s' "$SUBMIT_RESPONSE" | jq 'if has("data") then .data else . end')
PREDICTION_ID=$(printf '%s' "$TASK" | jq -r '.id')
if [ -z "$PREDICTION_ID" ] || [ "$PREDICTION_ID" = "null" ]; then
printf 'Submission response did not contain a prediction id
' >&2
exit 1
fi
RESULT_URL="https://api.wavespeed.ai/api/v3/predictions/$PREDICTION_ID/result"
# 2. Poll until the prediction finishes.
while true; do
RESPONSE=$(curl --silent --show-error --fail-with-body "$RESULT_URL" \
-H "Authorization: Bearer $WAVESPEED_API_KEY")
RESULT=$(printf '%s' "$RESPONSE" | jq 'if has("data") then .data else . end')
STATUS=$(printf '%s' "$RESULT" | jq -r '.status')
case "$STATUS" in
completed) printf '%s\n' "$RESULT" | jq '.outputs'; break ;;
failed|cancelled|timeout|deleted) printf '%s\n' "$RESULT" | jq . >&2; exit 1 ;;
*) sleep 2 ;;
esac
doneconst submitUrl = "https://api.wavespeed.ai/api/v3/wavespeed-ai/sam3-video-rle";
const apiKey = process.env.WAVESPEED_API_KEY;
if (!apiKey) throw new Error('Set WAVESPEED_API_KEY');
async function requestJson(url, options = {}) {
const response = await fetch(url, options);
if (!response.ok) throw new Error(await response.text());
return response.json();
}
// 1. Submit the prediction.
const body = await requestJson(submitUrl, {
method: "POST",
headers: {
"Authorization": `Bearer ${apiKey}`,
"Content-Type": "application/json",
},
body: JSON.stringify({
"prompt": "A cinematic shot of a city at sunset, soft golden light",
"video": "https://interactive-examples.mdn.mozilla.net/media/cc0-videos/flower.mp4",
"apply_mask": true
}),
});
const task = body.data ?? body;
if (!task.id) throw new Error("Submission response did not contain a prediction id");
const resultUrl = `https://api.wavespeed.ai/api/v3/predictions/${task.id}/result`;
// 2. Poll until the prediction finishes.
while (true) {
const resultBody = await requestJson(resultUrl, {
headers: { "Authorization": `Bearer ${apiKey}` },
});
const result = resultBody.data ?? resultBody;
if (result.status === "completed") {
console.log(result.outputs);
break;
}
if (["failed", "cancelled", "timeout", "deleted"].includes(result.status)) throw new Error(JSON.stringify(result));
await new Promise(resolve => setTimeout(resolve, 2000));
}import json
import os
import time
from urllib.request import Request, urlopen
api_key = os.environ["WAVESPEED_API_KEY"]
headers = {"Authorization": f"Bearer {api_key}", "Content-Type": "application/json"}
payload = {
"prompt": "A cinematic shot of a city at sunset, soft golden light",
"video": "https://interactive-examples.mdn.mozilla.net/media/cc0-videos/flower.mp4",
"apply_mask": True
}
def request_json(url, data=None):
request = Request(url, data=data, headers=headers, method="POST" if data else "GET")
with urlopen(request) as response:
return json.load(response)
# 1. Submit the prediction.
body = request_json("https://api.wavespeed.ai/api/v3/wavespeed-ai/sam3-video-rle", json.dumps(payload).encode())
task = body.get("data", body)
if not task.get("id"):
raise RuntimeError("Submission response did not contain a prediction id")
result_url = f"https://api.wavespeed.ai/api/v3/predictions/{task['id']}/result"
# 2. Poll until the prediction finishes.
while True:
result_body = request_json(result_url)
result = result_body.get("data", result_body)
status = result.get("status")
if status == "completed":
print(result.get("outputs", []))
break
if status in {"failed", "cancelled", "timeout", "deleted"}:
raise RuntimeError(result)
time.sleep(2)Sam3 Video Rle is a WaveSpeedAI model for AI inference, exposed as a REST API on WaveSpeedAI. SAM 3 Video RLE is a unified foundation model for prompt-based segmentation in video. Track and segment objects across frames using text, points, or boxes, returning RLE encoded masks for efficient processing. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing. You can call it programmatically or try it from the playground above.
POST your input parameters to the model's REST endpoint (shown in the API tab of this playground) with your WaveSpeedAI API key in the Authorization header. Submission returns a prediction ID. Poll the result endpoint starting around every 2 seconds, increase the interval for long-running tasks, and stop on any terminal status. The playground generates production-oriented Python, JavaScript, and cURL examples with timeouts, transient-error handling, and safe GET retries. Full request/response shape is documented at https://wavespeed.ai/docs/docs-api/wavespeed-ai/sam3-video-rle.
Sam3 Video Rle starts at $0.050 per run. That figure is the base price — the final charge scales with the parameters you set in the form (output size, length, count, references, or whatever knobs this model exposes), so a higher-quality or larger output costs more than a minimal one. The exact cost for your current input is shown live next to the Generate button before you submit, and the actual per-call charge is recorded on the prediction afterwards.
Key inputs: `prompt`, `video`, `apply_mask`, `box_prompts`, `point_prompts`. The full JSON schema (types, defaults, allowed values) is rendered above the Generate button and mirrored in the API reference at https://wavespeed.ai/docs/docs-api/wavespeed-ai/sam3-video-rle.
Median end-to-end generation time on WaveSpeedAI is around 74 seconds per request, based on recent successful runs. Queue time varies with global demand; live status is visible in the prediction record.
Commercial usage rights depend on the model's license, set by its provider (WaveSpeedAI). The license summary appears on the model card above; see WaveSpeedAI's Terms of Service for platform-level conditions.