Seedance 2.0 20 % DE DESCUENTO | Crea en el Video Generator →

Soulx Flashhead

wavespeed-ai /

SoulX FlashHead enables real-time streaming talking head video generation from portrait image and audio with ultra-fast 96 FPS performance. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

digital-human
Entrada

Arrastra y suelta o haz clic para subir

preview

Arrastra y suelta o haz clic para subir

Inactivo

$0.075por ejecución·~13 / $1

EjemplosVer todo

Modelos relacionados

README

SoulX FlashHead

SoulX FlashHead generates realistic talking head videos by combining a portrait image with audio. Upload a face image and audio clip — the model creates a lip-synced video with natural facial movements and expressions. With support for audio clips up to 30 minutes and budget-friendly pricing, it's ideal for long-form talking avatar content.

Why Choose This?

  • Long-form support Generate talking avatar videos with audio up to 30 minutes in length.

  • Realistic lip-sync Accurate lip movements synchronized to the audio input.

  • Natural expressions Creates realistic facial movements and expressions during speech.

  • Budget-friendly Lower cost per second compared to other talking avatar models.

  • Resolution options Choose between 480p and 720p output quality.

Parameters

ParameterRequiredDescription
imageYesPortrait image for the avatar (URL or upload)
audioYesAudio clip for lip-sync (URL or upload, max: 30 min)
resolutionNoOutput resolution: 480p, 720p (default)
seedNoRandom seed for reproducibility (-1 for random)

How to Use

  1. Upload your image — provide a clear portrait image with visible face.
  2. Upload your audio — add the audio clip you want the avatar to speak (up to 30 minutes).
  3. Select resolution — choose 720p for higher quality or 480p for faster/cheaper processing.
  4. Run — submit and download your talking avatar video.

Pricing

Duration720p480p
≤5 s$0.15$0.075
10 s$0.30$0.15
30 s$0.90$0.45
60 s$1.80$0.90
5 min$9.00$4.50

Billing Rules

  • Minimum charge: 5 seconds
  • Maximum duration: 1800 seconds (30 minutes)
  • 480p rate: $0.075 per 5 seconds
  • 720p rate: $0.15 per 5 seconds (2× 480p rate)

Best Use Cases

  • Long-Form Content — Create extended talking avatar videos for podcasts, lectures, and presentations.
  • E-learning — Generate instructor avatars for educational courses and tutorials.
  • Digital Presenters — Produce talking avatars for video presentations and explainers.
  • Marketing & Ads — Create spokesperson videos without filming.
  • Localization — Generate lip-synced videos in different languages from the same avatar.

Pro Tips

  • Use clear, front-facing portrait images with good lighting for best results.
  • Ensure the face is clearly visible and not obscured by hair or accessories.
  • Use high-quality audio with clear speech for accurate lip-sync.
  • Choose 480p for longer content to reduce costs — it still provides good quality.
  • Set a specific seed for reproducible results across multiple generations.

Notes

  • Both image and audio are required fields.
  • Maximum audio duration: 30 minutes (1800 seconds).
  • Minimum charge: 5 seconds (shorter clips billed as 5 seconds).
  • 720p resolution costs 2× the 480p rate.
  • Ensure uploaded file URLs are publicly accessible.

Related Models

Accesibilidad:Este sitio web utiliza modelos de IA proporcionados por terceros.

Soulx Flashhead API — Quick start

Grab a WaveSpeedAI API key, then call POST https://api.wavespeed.ai/api/v3/wavespeed-ai/soulx-flashhead with your input as JSON. The endpoint returns a prediction id; poll the prediction endpoint until status flips to completed, then read the output URL from data.outputs[0]. Examples for Soulx Flashhead below.

HTTP example
# Submit the prediction
curl -X POST "https://api.wavespeed.ai/api/v3/wavespeed-ai/soulx-flashhead" \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $WAVESPEED_API_KEY" \
  -d '{
    "image": "https://example.com/your-input.jpg",
    "audio": "https://example.com/your-audio.mp3",
    "resolution": "720p",
    "seed": -1
}'

# Response includes a prediction id. Poll for the result:
curl -X GET "https://api.wavespeed.ai/api/v3/predictions/{request_id}/result" \
  -H "Authorization: Bearer $WAVESPEED_API_KEY"

# When status is "completed", read the output from data.outputs[0].
Node.js example
// npm install wavespeed
const WaveSpeed = require('wavespeed');

const client = new WaveSpeed(); // reads WAVESPEED_API_KEY from env

const result = await client.run("wavespeed-ai/soulx-flashhead", {
        "image": "https://example.com/your-input.jpg",
        "audio": "https://example.com/your-audio.mp3",
        "resolution": "720p",
        "seed": -1
});

console.log(result.outputs[0]); // → URL of the generated output
Python example
# pip install wavespeed
import wavespeed

output = wavespeed.run(
    "wavespeed-ai/soulx-flashhead",
    {
    "image": "https://example.com/your-input.jpg",
    "audio": "https://example.com/your-audio.mp3",
    "resolution": "720p",
    "seed": -1
}
)

print(output["outputs"][0])  # → URL of the generated output

Soulx Flashhead API — Frequently asked questions

What is the Soulx Flashhead API?

Soulx Flashhead is a WaveSpeedAI model for talking-avatar generation, exposed as a REST API on WaveSpeedAI. SoulX FlashHead enables real-time streaming talking head video generation from portrait image and audio with ultra-fast 96 FPS performance. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing. You can call it programmatically or try it from the playground above.

How do I call the Soulx Flashhead API?

POST your input parameters to the model's REST endpoint (shown in the API tab of this playground) with your WaveSpeedAI API key in the Authorization header. Submission returns a prediction ID; poll the prediction endpoint until status flips to "completed", then read the output URL from the result. The playground generates a ready-to-paste code sample in Python, JavaScript, or cURL for whatever inputs you've set. Full request/response shape is documented at https://wavespeed.ai/docs/docs-api/wavespeed-ai/soulx-flashhead.

How much does Soulx Flashhead cost per run?

Soulx Flashhead starts at $0.075 per run. That figure is the base price — the final charge scales with the parameters you set in the form (output size, length, count, references, or whatever knobs this model exposes), so a higher-quality or larger output costs more than a minimal one. The exact cost for your current input is shown live next to the Generate button before you submit, and the actual per-call charge is recorded on the prediction afterwards.

What inputs does Soulx Flashhead accept?

Key inputs: `image`, `audio`, `resolution`, `seed`. The full JSON schema (types, defaults, allowed values) is rendered above the Generate button and mirrored in the API reference at https://wavespeed.ai/docs/docs-api/wavespeed-ai/soulx-flashhead.

How do I get started with the Soulx Flashhead API?

Sign up for a free WaveSpeedAI account to claim starter credits, copy your API key from /accesskey, then call the endpoint shown in the API tab of the playground. The playground also auto-generates a code sample in Python, JavaScript, or cURL for the parameters you've set.

Can I use Soulx Flashhead outputs commercially?

Commercial usage rights depend on the model's license, set by its provider (WaveSpeedAI). The license summary appears on the model card above; see WaveSpeedAI's Terms of Service for platform-level conditions.