Veed Clean Audio API Documentation
Playground
Try it on WaveSpeedAI!VEED Clean Audio cleans speech recordings with adjustable noise suppression and loudness control, improving voice clarity for podcasts, interviews, voiceovers, meetings, and audio production workflows. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
Features
VEED Clean Audio improves speech recordings by reducing background noise and normalizing loudness for clearer, more consistent spoken audio. Upload an audio file or a video containing an audio track, adjust the suppression strength and target loudness if needed, and export the cleaned result as FLAC or WAV.
It is designed for interviews, podcasts, voiceovers, meetings, presentations, and other speech-focused recordings where background noise or uneven loudness reduces clarity.
Why Choose This?
-
Speech-focused cleanup
Reduce distracting background noise while keeping spoken content clear. -
Adjustable noise suppression
Usestrengthto control how aggressively background noise and room tone are reduced. -
Loudness normalization
Set a target loudness withtarget_lufsfor more consistent playback levels. -
Audio from video input
Process the audio track from a video file and return cleaned audio. -
Lossless output options
Export the processed result asflacorwav. -
Up to 30 minutes
Process recordings up to30minutes per request.
Parameters
| Parameter | Required | Description |
|---|---|---|
| audio | Yes | Public URL of an audio recording or video containing an audio track. Maximum duration: 30 minutes. Maximum file size: 512 MB. |
| strength | No | Noise suppression strength from 0 to 1. Default: 0.874. |
| target_lufs | No | Target output loudness from -40 to -8 LUFS. Default: -19. |
| output_format | No | Output audio format: flac or wav. Default: flac. |
How to Use
- Provide a recording — Upload an audio file or a video containing an audio track.
- Adjust suppression optional — Change
strengthto control how much background noise is removed. - Set loudness optional — Adjust
target_lufsor keep the default-19. - Choose output format — Select
flacorwav. - Submit — Clean the recording and retrieve the processed audio.
Pricing
Pricing is $0.014 per billable minute.
Input duration is rounded to the nearest whole second and then rounded up to the next full minute, with a minimum billed duration of 1 minute.
| Input Duration | Billable Minutes | Price |
|---|---|---|
| 30 seconds | 1 | $0.014 |
| 60 seconds | 1 | $0.014 |
| 61 seconds | 2 | $0.028 |
| 5 minutes | 5 | $0.070 |
| 10 minutes | 10 | $0.140 |
| 30 minutes | 30 | $0.420 |
The minimum charge is $0.014 per request.
strength, target_lufs, and output_format do not add separate charges.
Supported input duration is up to 30 minutes. Over-limit inputs are rejected rather than automatically truncated.
Best Use Cases
- Interviews and podcasts — Reduce environmental noise and create more consistent spoken audio.
- Voiceovers — Clean narration recorded outside a controlled studio environment.
- Meeting recordings — Improve speech clarity in conference, remote-call, or presentation recordings.
- Presentation audio — Normalize spoken content for more consistent playback.
- Video speech cleanup — Extract and clean the audio track from a video.
- Creator content — Prepare spoken audio for videos, podcasts, courses, and social media.
Pro Tips
- Start with the default
strengthbefore increasing noise suppression. - Lower
strengthwhen you want to preserve more natural room ambience. - Use a stronger setting when steady background noise is especially distracting.
- Keep
target_lufsclose to your delivery platform’s preferred loudness when consistency matters. - Use
wavfor broad editing compatibility. - Use
flacwhen you want lossless audio with a smaller file size.
Notes
audiois required.- Maximum input duration is
30minutes. - Maximum input size is
512 MB. - The input must contain decodable audio.
- File URLs must be public and directly accessible.
- Redirects are not followed.
- Multichannel audio is mixed down to mono.
- Video inputs return cleaned audio only.
- This endpoint cleans existing speech; it does not clone voices or generate new speech.
- Severe clipping, overlapping speakers, or very faint speech may limit cleanup quality.
Related Models
- VEED Subtitles — Generate and add subtitles for video content.
- VEED Fabric 1.0 — Create and edit video content with VEED’s Fabric workflow.
Authentication
For authentication details, please refer to the Authentication Guide.
API Endpoints
Submit Task & Query Result
set -euo pipefail
export WAVESPEED_API_KEY="your-api-key"
REQUEST_BODY=$(cat <<'JSON'
{
"audio": "https://interactive-examples.mdn.mozilla.net/media/cc0-videos/flower.mp4",
"strength": 0.874,
"target_lufs": -19,
"output_format": "flac"
}
JSON
)
# 1. Submit the prediction.
SUBMIT_RESPONSE=$(curl --silent --show-error --fail-with-body \
-X POST "https://api.wavespeed.ai/api/v3/veed/clean-audio" \
-H "Authorization: Bearer ${WAVESPEED_API_KEY}" \
-H "Content-Type: application/json" \
-d "${REQUEST_BODY}")
TASK=$(printf '%s' "${SUBMIT_RESPONSE}" | jq 'if type == "object" and has("data") then .data else . end')
PREDICTION_ID=$(printf '%s' "${TASK}" | jq -r '.id // empty')
if [ -z "${PREDICTION_ID}" ]; then
printf 'Submission response did not contain a prediction id
' >&2
exit 1
fi
RESULT_URL="https://api.wavespeed.ai/api/v3/predictions/${PREDICTION_ID}/result"
# 2. Poll until the prediction finishes.
while true; do
RESPONSE=$(curl --silent --show-error --fail-with-body \
"${RESULT_URL}" \
-H "Authorization: Bearer ${WAVESPEED_API_KEY}")
RESULT=$(printf '%s' "${RESPONSE}" | jq 'if type == "object" and has("data") then .data else . end')
STATUS=$(printf '%s' "${RESULT}" | jq -r '.status // empty')
case "${STATUS}" in
completed) printf '%s\n' "${RESULT}" | jq '.outputs'; break ;;
failed|cancelled|timeout|deleted) printf '%s\n' "${RESULT}" | jq . >&2; exit 1 ;;
*) sleep 2 ;;
esac
doneParameters
Task Submission Parameters
Request Parameters
| Parameter | Type | Required | Default | Range | Description |
|---|---|---|---|---|---|
| audio | string | Yes | - | - | Recording to clean. Accepts audio or a video's audio track. Maximum duration: 30 minutes; maximum file size: 512 MB. Use a public, directly accessible URL. |
| strength | number | No | 0.874 | 0 ~ 1 | Noise suppression strength. Lower values retain more room tone under speech. |
| target_lufs | number | No | -19 | -40 ~ -8 | Target integrated loudness in LUFS. True peak is limited to -1.1 dBTP. |
| output_format | string | No | flac | flac, wav | Lossless output container. Both formats use 48 kHz mono 16-bit audio. |
Response Parameters
| Parameter | Type | Description |
|---|---|---|
| code | integer | HTTP status code (e.g., 200 for success) |
| message | string | Status message (e.g., “success”) |
| data.id | string | Unique identifier for the prediction, Task Id |
| data.model | string | Model ID used for the prediction |
| data.outputs | array | Output values, usually URL strings; some models return text strings or structured result objects (empty when status is not completed) |
| data.urls | object | Object containing related API endpoints |
| data.status | string | Task status. completed is successful; failed, cancelled, timeout, and deleted are failure terminal statuses. |
| data.created_at | string | ISO timestamp of when the request was created (e.g., “2023-04-01T12:34:56.789Z”) |
| data.error | string | Error message (empty if no error occurred) |
| data.timings | object | Object containing timing details |
| data.timings.inference | integer | Inference time in milliseconds |
Result Request Parameters
| Parameter | Type | Required | Default | Description |
|---|---|---|---|---|
| id | string | Yes | - | Task ID |
Result Response Parameters
| Parameter | Type | Description |
|---|---|---|
| code | integer | HTTP status code (e.g., 200 for success) |
| message | string | Status message (e.g., “success”) |
| data | object | The prediction data object containing all details |
| data.id | string | Unique identifier for the prediction |
| data.model | string | Model ID used for the prediction |
| data.outputs | array<string | object> | Array of generated outputs (empty when status is not completed). Items are usually URL strings, but may be text strings or structured result objects, depending on the model. |
| data.urls | object | Object containing related API endpoints |
| data.status | string | Status: completed is successful; failed, cancelled, timeout, and deleted are failure terminal statuses |
| data.created_at | string | ISO timestamp of when the request was created |
| data.error | string | Error message (empty if no error occurred) |
| data.timings | object | Object containing timing details |
| data.timings.inference | integer | Inference time in milliseconds |