Seedream 5.0 Flash is LIVE — Faster & Cheaper | Try Now →

veed/clean-audio

VEED Clean Audio cleans speech recordings with adjustable noise suppression and loudness control, improving voice clarity for podcasts, interviews, voiceovers, meetings, and audio production workflows. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

audio-to-audio
Input
Enable Safety Checker

Idle

$0.014per run·~71 / $1

ExamplesView all

Related Models

README

VEED Clean Audio

VEED Clean Audio improves speech recordings by reducing background noise and normalizing loudness for clearer, more consistent spoken audio. Upload an audio file or a video containing an audio track, adjust the suppression strength and target loudness if needed, and export the cleaned result as FLAC or WAV.

It is designed for interviews, podcasts, voiceovers, meetings, presentations, and other speech-focused recordings where background noise or uneven loudness reduces clarity.

Why Choose This?

  • Speech-focused cleanup
    Reduce distracting background noise while keeping spoken content clear.

  • Adjustable noise suppression
    Use strength to control how aggressively background noise and room tone are reduced.

  • Loudness normalization
    Set a target loudness with target_lufs for more consistent playback levels.

  • Audio from video input
    Process the audio track from a video file and return cleaned audio.

  • Lossless output options
    Export the processed result as flac or wav.

  • Up to 30 minutes
    Process recordings up to 30 minutes per request.

Parameters

ParameterRequiredDescription
audioYesPublic URL of an audio recording or video containing an audio track. Maximum duration: 30 minutes. Maximum file size: 512 MB.
strengthNoNoise suppression strength from 0 to 1. Default: 0.874.
target_lufsNoTarget output loudness from -40 to -8 LUFS. Default: -19.
output_formatNoOutput audio format: flac or wav. Default: flac.

How to Use

  1. Provide a recording — Upload an audio file or a video containing an audio track.
  2. Adjust suppression optional — Change strength to control how much background noise is removed.
  3. Set loudness optional — Adjust target_lufs or keep the default -19.
  4. Choose output format — Select flac or wav.
  5. Submit — Clean the recording and retrieve the processed audio.

Pricing

Pricing is $0.014 per billable minute.

Input duration is rounded to the nearest whole second and then rounded up to the next full minute, with a minimum billed duration of 1 minute.

Input DurationBillable MinutesPrice
30 seconds1$0.014
60 seconds1$0.014
61 seconds2$0.028
5 minutes5$0.070
10 minutes10$0.140
30 minutes30$0.420

The minimum charge is $0.014 per request.

strength, target_lufs, and output_format do not add separate charges.

Supported input duration is up to 30 minutes. Over-limit inputs are rejected rather than automatically truncated.

Best Use Cases

  • Interviews and podcasts — Reduce environmental noise and create more consistent spoken audio.
  • Voiceovers — Clean narration recorded outside a controlled studio environment.
  • Meeting recordings — Improve speech clarity in conference, remote-call, or presentation recordings.
  • Presentation audio — Normalize spoken content for more consistent playback.
  • Video speech cleanup — Extract and clean the audio track from a video.
  • Creator content — Prepare spoken audio for videos, podcasts, courses, and social media.

Pro Tips

  • Start with the default strength before increasing noise suppression.
  • Lower strength when you want to preserve more natural room ambience.
  • Use a stronger setting when steady background noise is especially distracting.
  • Keep target_lufs close to your delivery platform's preferred loudness when consistency matters.
  • Use wav for broad editing compatibility.
  • Use flac when you want lossless audio with a smaller file size.

Notes

  • audio is required.
  • Maximum input duration is 30 minutes.
  • Maximum input size is 512 MB.
  • The input must contain decodable audio.
  • File URLs must be public and directly accessible.
  • Redirects are not followed.
  • Multichannel audio is mixed down to mono.
  • Video inputs return cleaned audio only.
  • This endpoint cleans existing speech; it does not clone voices or generate new speech.
  • Severe clipping, overlapping speakers, or very faint speech may limit cleanup quality.

Related Models

  • VEED Subtitles — Generate and add subtitles for video content.
  • VEED Fabric 1.0 — Create and edit video content with VEED's Fabric workflow.
Note:This website uses AI models provided by third parties. Documentation prices are for reference and may be outdated. The Generate button shows an estimate; the final task charge prevails.

Clean Audio API — Quick start

Grab a WaveSpeedAI API key, then call POST https://api.wavespeed.ai/api/v3/veed/clean-audio with your input as JSON. The endpoint returns a prediction id. Start polling the result endpoint around every 2 seconds, increase the interval for long-running tasks, and stop on any terminal status. On completed, read output values from data.outputs. Examples for Clean Audio below.

HTTP example
set -euo pipefail

: "${WAVESPEED_API_KEY:?Set WAVESPEED_API_KEY}"

REQUEST_BODY=$(cat <<'JSON'
{
    "audio": "https://interactive-examples.mdn.mozilla.net/media/cc0-audio/t-rex-roar.mp3",
    "strength": 0.874,
    "target_lufs": -19,
    "output_format": "flac"
}
JSON
)

# 1. Submit the prediction.
SUBMIT_RESPONSE=$(curl --silent --show-error --fail-with-body \
  -X POST "https://api.wavespeed.ai/api/v3/veed/clean-audio" \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $WAVESPEED_API_KEY" \
  -d "$REQUEST_BODY")

TASK=$(printf '%s' "$SUBMIT_RESPONSE" | jq 'if has("data") then .data else . end')
PREDICTION_ID=$(printf '%s' "$TASK" | jq -r '.id')
if [ -z "$PREDICTION_ID" ] || [ "$PREDICTION_ID" = "null" ]; then
  printf 'Submission response did not contain a prediction id
' >&2
  exit 1
fi
RESULT_URL="https://api.wavespeed.ai/api/v3/predictions/$PREDICTION_ID/result"

# 2. Poll until the prediction finishes.
while true; do
  RESPONSE=$(curl --silent --show-error --fail-with-body "$RESULT_URL" \
    -H "Authorization: Bearer $WAVESPEED_API_KEY")
  RESULT=$(printf '%s' "$RESPONSE" | jq 'if has("data") then .data else . end')
  STATUS=$(printf '%s' "$RESULT" | jq -r '.status')
  case "$STATUS" in
    completed) printf '%s\n' "$RESULT" | jq '.outputs'; break ;;
    failed|cancelled|timeout|deleted) printf '%s\n' "$RESULT" | jq . >&2; exit 1 ;;
    *) sleep 2 ;;
  esac
done
Node.js example
const submitUrl = "https://api.wavespeed.ai/api/v3/veed/clean-audio";
const apiKey = process.env.WAVESPEED_API_KEY;
if (!apiKey) throw new Error('Set WAVESPEED_API_KEY');

async function requestJson(url, options = {}) {
  const response = await fetch(url, options);
  if (!response.ok) throw new Error(await response.text());
  return response.json();
}

// 1. Submit the prediction.
const body = await requestJson(submitUrl, {
  method: "POST",
  headers: {
    "Authorization": `Bearer ${apiKey}`,
    "Content-Type": "application/json",
  },
  body: JSON.stringify({
        "audio": "https://interactive-examples.mdn.mozilla.net/media/cc0-audio/t-rex-roar.mp3",
        "strength": 0.874,
        "target_lufs": -19,
        "output_format": "flac"
}),
});
const task = body.data ?? body;
if (!task.id) throw new Error("Submission response did not contain a prediction id");
const resultUrl = `https://api.wavespeed.ai/api/v3/predictions/${task.id}/result`;

// 2. Poll until the prediction finishes.
while (true) {
  const resultBody = await requestJson(resultUrl, {
    headers: { "Authorization": `Bearer ${apiKey}` },
  });
  const result = resultBody.data ?? resultBody;
  if (result.status === "completed") {
    console.log(result.outputs);
    break;
  }
  if (["failed", "cancelled", "timeout", "deleted"].includes(result.status)) throw new Error(JSON.stringify(result));
  await new Promise(resolve => setTimeout(resolve, 2000));
}
Python example
import json
import os
import time
from urllib.request import Request, urlopen

api_key = os.environ["WAVESPEED_API_KEY"]
headers = {"Authorization": f"Bearer {api_key}", "Content-Type": "application/json"}
payload = {
    "audio": "https://interactive-examples.mdn.mozilla.net/media/cc0-audio/t-rex-roar.mp3",
    "strength": 0.874,
    "target_lufs": -19,
    "output_format": "flac"
}

def request_json(url, data=None):
    request = Request(url, data=data, headers=headers, method="POST" if data else "GET")
    with urlopen(request) as response:
        return json.load(response)

# 1. Submit the prediction.
body = request_json("https://api.wavespeed.ai/api/v3/veed/clean-audio", json.dumps(payload).encode())
task = body.get("data", body)
if not task.get("id"):
    raise RuntimeError("Submission response did not contain a prediction id")
result_url = f"https://api.wavespeed.ai/api/v3/predictions/{task['id']}/result"

# 2. Poll until the prediction finishes.
while True:
    result_body = request_json(result_url)
    result = result_body.get("data", result_body)
    status = result.get("status")
    if status == "completed":
        print(result.get("outputs", []))
        break
    if status in {"failed", "cancelled", "timeout", "deleted"}:
        raise RuntimeError(result)
    time.sleep(2)

Clean Audio API — Frequently asked questions

What is the Clean Audio API?

Clean Audio is a Veed model for AI inference, exposed as a REST API on WaveSpeedAI. VEED Clean Audio cleans speech recordings with adjustable noise suppression and loudness control, improving voice clarity for podcasts, interviews, voiceovers, meetings, and audio production workflows. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing. You can call it programmatically or try it from the playground above.

How do I call the Clean Audio API?

POST your input parameters to the model's REST endpoint (shown in the API tab of this playground) with your WaveSpeedAI API key in the Authorization header. Submission returns a prediction ID. Poll the result endpoint starting around every 2 seconds, increase the interval for long-running tasks, and stop on any terminal status. The playground generates Python, JavaScript, and cURL examples for submitting requests and polling results. Full request/response shape is documented at https://wavespeed.ai/docs/docs-api/veed/veed-clean-audio.

How much does Clean Audio cost per run?

Clean Audio starts at $0.014 per run. That figure is the base price — the final charge scales with the parameters you set in the form (output size, length, count, references, or whatever knobs this model exposes), so a higher-quality or larger output costs more than a minimal one. The exact cost for your current input is shown live next to the Generate button before you submit, and the actual per-call charge is recorded on the prediction afterwards.

What inputs does Clean Audio accept?

Key inputs: `audio`, `output_format`, `strength`, `target_lufs`. The full JSON schema (types, defaults, allowed values) is rendered above the Generate button and mirrored in the API reference at https://wavespeed.ai/docs/docs-api/veed/veed-clean-audio.

How do I get started with the Clean Audio API?

Sign up for a free WaveSpeedAI account to claim starter credits, copy your API key from /accesskey, then call the endpoint shown in the API tab of the playground. The playground also auto-generates a code sample in Python, JavaScript, or cURL for the parameters you've set.

Can I use Clean Audio outputs commercially?

Commercial usage rights depend on the model's license, set by its provider (Veed). Check the provider's applicable terms and WaveSpeedAI's Terms of Service before commercial use.

llms.txt — veed/clean-audio for AI agents and LLMs