Veed Clean Audio API Documentation

Veed Clean Audio API Documentation

Playground

Try it on WaveSpeedAI!

VEED Clean Audio cleans speech recordings with adjustable noise suppression and loudness control, improving voice clarity for podcasts, interviews, voiceovers, meetings, and audio production workflows. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

Features

VEED Clean Audio improves speech recordings by reducing background noise and normalizing loudness for clearer, more consistent spoken audio. Upload an audio file or a video containing an audio track, adjust the suppression strength and target loudness if needed, and export the cleaned result as FLAC or WAV.

It is designed for interviews, podcasts, voiceovers, meetings, presentations, and other speech-focused recordings where background noise or uneven loudness reduces clarity.


Why Choose This?

  • Speech-focused cleanup
    Reduce distracting background noise while keeping spoken content clear.

  • Adjustable noise suppression
    Use strength to control how aggressively background noise and room tone are reduced.

  • Loudness normalization
    Set a target loudness with target_lufs for more consistent playback levels.

  • Audio from video input
    Process the audio track from a video file and return cleaned audio.

  • Lossless output options
    Export the processed result as flac or wav.

  • Up to 30 minutes
    Process recordings up to 30 minutes per request.


Parameters

ParameterRequiredDescription
audioYesPublic URL of an audio recording or video containing an audio track. Maximum duration: 30 minutes. Maximum file size: 512 MB.
strengthNoNoise suppression strength from 0 to 1. Default: 0.874.
target_lufsNoTarget output loudness from -40 to -8 LUFS. Default: -19.
output_formatNoOutput audio format: flac or wav. Default: flac.

How to Use

  1. Provide a recording — Upload an audio file or a video containing an audio track.
  2. Adjust suppression optional — Change strength to control how much background noise is removed.
  3. Set loudness optional — Adjust target_lufs or keep the default -19.
  4. Choose output format — Select flac or wav.
  5. Submit — Clean the recording and retrieve the processed audio.

Pricing

Pricing is $0.014 per billable minute.

Input duration is rounded to the nearest whole second and then rounded up to the next full minute, with a minimum billed duration of 1 minute.

Input DurationBillable MinutesPrice
30 seconds1$0.014
60 seconds1$0.014
61 seconds2$0.028
5 minutes5$0.070
10 minutes10$0.140
30 minutes30$0.420

The minimum charge is $0.014 per request.

strength, target_lufs, and output_format do not add separate charges.

Supported input duration is up to 30 minutes. Over-limit inputs are rejected rather than automatically truncated.


Best Use Cases

  • Interviews and podcasts — Reduce environmental noise and create more consistent spoken audio.
  • Voiceovers — Clean narration recorded outside a controlled studio environment.
  • Meeting recordings — Improve speech clarity in conference, remote-call, or presentation recordings.
  • Presentation audio — Normalize spoken content for more consistent playback.
  • Video speech cleanup — Extract and clean the audio track from a video.
  • Creator content — Prepare spoken audio for videos, podcasts, courses, and social media.

Pro Tips

  • Start with the default strength before increasing noise suppression.
  • Lower strength when you want to preserve more natural room ambience.
  • Use a stronger setting when steady background noise is especially distracting.
  • Keep target_lufs close to your delivery platform’s preferred loudness when consistency matters.
  • Use wav for broad editing compatibility.
  • Use flac when you want lossless audio with a smaller file size.

Notes

  • audio is required.
  • Maximum input duration is 30 minutes.
  • Maximum input size is 512 MB.
  • The input must contain decodable audio.
  • File URLs must be public and directly accessible.
  • Redirects are not followed.
  • Multichannel audio is mixed down to mono.
  • Video inputs return cleaned audio only.
  • This endpoint cleans existing speech; it does not clone voices or generate new speech.
  • Severe clipping, overlapping speakers, or very faint speech may limit cleanup quality.

  • VEED Subtitles — Generate and add subtitles for video content.
  • VEED Fabric 1.0 — Create and edit video content with VEED’s Fabric workflow.

Authentication

For authentication details, please refer to the Authentication Guide.

API Endpoints

Submit Task & Query Result

set -euo pipefail

export WAVESPEED_API_KEY="your-api-key"

REQUEST_BODY=$(cat <<'JSON'
{
  "audio": "https://interactive-examples.mdn.mozilla.net/media/cc0-videos/flower.mp4",
  "strength": 0.874,
  "target_lufs": -19,
  "output_format": "flac"
}
JSON
)

# 1. Submit the prediction.
SUBMIT_RESPONSE=$(curl --silent --show-error --fail-with-body \
  -X POST "https://api.wavespeed.ai/api/v3/veed/clean-audio" \
  -H "Authorization: Bearer ${WAVESPEED_API_KEY}" \
  -H "Content-Type: application/json" \
  -d "${REQUEST_BODY}")

TASK=$(printf '%s' "${SUBMIT_RESPONSE}" | jq 'if type == "object" and has("data") then .data else . end')
PREDICTION_ID=$(printf '%s' "${TASK}" | jq -r '.id // empty')
if [ -z "${PREDICTION_ID}" ]; then
  printf 'Submission response did not contain a prediction id
' >&2
  exit 1
fi
RESULT_URL="https://api.wavespeed.ai/api/v3/predictions/${PREDICTION_ID}/result"

# 2. Poll until the prediction finishes.
while true; do
  RESPONSE=$(curl --silent --show-error --fail-with-body \
    "${RESULT_URL}" \
    -H "Authorization: Bearer ${WAVESPEED_API_KEY}")
  RESULT=$(printf '%s' "${RESPONSE}" | jq 'if type == "object" and has("data") then .data else . end')
  STATUS=$(printf '%s' "${RESULT}" | jq -r '.status // empty')

  case "${STATUS}" in
    completed) printf '%s\n' "${RESULT}" | jq '.outputs'; break ;;
    failed|cancelled|timeout|deleted) printf '%s\n' "${RESULT}" | jq . >&2; exit 1 ;;
    *) sleep 2 ;;
  esac
done

Parameters

Task Submission Parameters

Request Parameters

ParameterTypeRequiredDefaultRangeDescription
audiostringYes--Recording to clean. Accepts audio or a video's audio track. Maximum duration: 30 minutes; maximum file size: 512 MB. Use a public, directly accessible URL.
strengthnumberNo0.8740 ~ 1Noise suppression strength. Lower values retain more room tone under speech.
target_lufsnumberNo-19-40 ~ -8Target integrated loudness in LUFS. True peak is limited to -1.1 dBTP.
output_formatstringNoflacflac, wavLossless output container. Both formats use 48 kHz mono 16-bit audio.

Response Parameters

ParameterTypeDescription
codeintegerHTTP status code (e.g., 200 for success)
messagestringStatus message (e.g., “success”)
data.idstringUnique identifier for the prediction, Task Id
data.modelstringModel ID used for the prediction
data.outputsarrayOutput values, usually URL strings; some models return text strings or structured result objects (empty when status is not completed)
data.urlsobjectObject containing related API endpoints
data.statusstringTask status. completed is successful; failed, cancelled, timeout, and deleted are failure terminal statuses.
data.created_atstringISO timestamp of when the request was created (e.g., “2023-04-01T12:34:56.789Z”)
data.errorstringError message (empty if no error occurred)
data.timingsobjectObject containing timing details
data.timings.inferenceintegerInference time in milliseconds

Result Request Parameters

ParameterTypeRequiredDefaultDescription
idstringYes-Task ID

Result Response Parameters

ParameterTypeDescription
codeintegerHTTP status code (e.g., 200 for success)
messagestringStatus message (e.g., “success”)
dataobjectThe prediction data object containing all details
data.idstringUnique identifier for the prediction
data.modelstringModel ID used for the prediction
data.outputsarray<string | object>Array of generated outputs (empty when status is not completed). Items are usually URL strings, but may be text strings or structured result objects, depending on the model.
data.urlsobjectObject containing related API endpoints
data.statusstringStatus: completed is successful; failed, cancelled, timeout, and deleted are failure terminal statuses
data.created_atstringISO timestamp of when the request was created
data.errorstringError message (empty if no error occurred)
data.timingsobjectObject containing timing details
data.timings.inferenceintegerInference time in milliseconds
© 2026 WaveSpeedAI. All rights reserved.