Vidu Q3 Pro ist online — jetzt ausprobieren
OpenAI Models

OpenAI Models

OpenAI's state-of-the-art AI models for text, image, and multimodal applications, Sora 2 is included

OpenAI's state-of-the-art AI models for text, image, and multimodal applications, Sora 2 is included

Alle Modelle

22 Modelle
openai/gpt-image-2/edit
image-to-image

openai/gpt-image-2/edit

OpenAI's GPT Image 2 Edit enables image editing from natural-language instructions with one or more reference images. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

openai/gpt-image-2/text-to-image
text-to-image

openai/gpt-image-2/text-to-image

OpenAI's GPT Image 2 Text-to-Image generates high-quality images from natural-language prompts. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

wavespeed-ai/openai-whisper
speech-to-text

wavespeed-ai/openai-whisper

Whisper Large v3 speech-to-text: instant, accurate multilingual transcripts with automatic language detection and punctuation. Upload audio to get transcripts. Ready-to-use REST API, no coldstarts, affordable pricing.

wavespeed-ai/openai-whisper-turbo
speech-to-text

wavespeed-ai/openai-whisper-turbo

Accurate speech-to-text with OpenAI Whisper Large v3 Turbo: multilingual transcripts with auto language detection and punctuation. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

openai/dall-e-2
text-to-image

openai/dall-e-2

Original DALL-E 2 from OpenAI for classic text-to-image generation via the OpenAI Image Generation API. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

openai/gpt-image-1-high-fidelity
text-to-image

openai/gpt-image-1-high-fidelity

OpenAI GPT Image 1 High-Fidelity produces photorealistic, high-detail images for creative and production workflows, delivering improved texture and color fidelity. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

openai/sora-2/image-to-video
image-to-video

openai/sora-2/image-to-video

OpenAI Sora 2 generates realistic image-to-video content with synchronized audio, improved physics, sharper realism and steerability. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

openai/sora-2/image-to-video-pro
image-to-video

openai/sora-2/image-to-video-pro

OpenAI Sora 2 Image-to-Video Pro creates physics-aware, realistic videos with synchronized audio and greater steerability. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

openai/sora-2/text-to-video
text-to-video

openai/sora-2/text-to-video

OpenAI Sora 2 is a state-of-the-art text-to-video model with realistic visuals, accurate physics, synchronized audio, and strong steerability. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

openai/sora-2/text-to-video-pro
text-to-video

openai/sora-2/text-to-video-pro

OpenAI Sora 2 Text-to-Video Pro creates high-fidelity videos with synchronized audio, realistic physics, and enhanced steerability. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

openai/sora
text-to-video

openai/sora

Sora is OpenAI's multi-modal model that generates videos from text, images, or existing video inputs. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

openai/dall-e-3
text-to-image

openai/dall-e-3

OpenAI DALL·E 3 for high-fidelity text-to-image generation available as a managed API on WaveSpeedAI. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

openai/gpt-image-1-mini/text-to-image
text-to-image

openai/gpt-image-1-mini/text-to-image

GPT Image 1 Mini is a cost-efficient multimodal OpenAI model powered by GPT-5 that turns text or image prompts into high-quality images. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

openai/gpt-image-1-mini/edit
image-to-image

openai/gpt-image-1-mini/edit

GPT Image 1 Mini is a cost-efficient, natively multimodal OpenAI model that pairs GPT-5 language understanding with compact image editing and generation from text and image inputs to produce high-quality images. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

openai/gpt-image-1.5/text-to-image
text-to-image

openai/gpt-image-1.5/text-to-image

GPT Image 1.5 text to image is OpenAI’s fast, cost-efficient text-to-image generator powered by GPT-5 guidance. Create photorealistic shots, product renders, concept art, and stylized graphics from natural-language prompts (optionally conditioned with an image). Supports custom aspect ratios, seeds, negative prompts, hex color hints, and style presets. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

openai/gpt-image-1.5/edit
image-to-image

openai/gpt-image-1.5/edit

GPT Image 1.5 Edit is OpenAI’s image model for precise, natural-language edits. Add/remove objects, swap backgrounds, retouch faces, adjust colors/lighting, edit text/graphics, crop/resize, and apply hex color control. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

wavespeed-ai/openai-whisper-with-video
speech-to-text

wavespeed-ai/openai-whisper-with-video

OpenAI Whisper Large v3 (Video-to-Text) delivers high-accuracy multilingual transcription directly from video files, with automatic language detection and optional timestamped, subtitle-ready segments. Built for stable production use with a ready-to-use REST API, fast response, no cold starts, and predictable pricing.

openai/sora-2-pro/text-to-video
text-to-video

openai/sora-2-pro/text-to-video

OpenAI Sora 2 Pro is a state-of-the-art text-to-video model with realistic physics, synchronized audio, and strong steerability. Supports multiple resolutions up to 1080p and durations up to 20 seconds. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

openai/sora-2-pro/image-to-video
image-to-video

openai/sora-2-pro/image-to-video

OpenAI Sora 2 Pro Image-to-Video creates physics-aware, realistic videos from reference images with synchronized audio and strong steerability. Supports 720p and 1080p resolutions with durations up to 20 seconds. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

openai/sora-2/characters
video-to-text

openai/sora-2/characters

OpenAI Sora 2 Characters creates reusable character IDs from video references for consistent character appearance across Sora 2 generations. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

openai/gpt-image-1/text-to-image
text-to-image

openai/gpt-image-1/text-to-image

OpenAI GPT Image-1 generates images from text prompts from OpenAI's latest text-to-image model, ideal for creating visual assets. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

openai/gpt-image-1
image-to-image

openai/gpt-image-1

OpenAI's gpt-image-1 enables image generation and image editing via OpenAI's image API, ideal for creating and refining images. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

OpenAI Models

Cutting-edge OpenAI models across text, image, and multimodal creation—curated in one place. These models sit at the front line of generative AI, combining strong reasoning, cinematic rendering, and reliable performance for real-world workflows.

Catalog

  1. Sora-2 / Image-to-Video — Add motion to a single image with physics-aware dynamics and stable identities.
  2. Sora-2 / Image-to-Video Pro — Higher fidelity and longer, smoother camera language for editorial or production shots.
  3. Sora-2 / Text-to-Video — Generate scenes directly from text prompts; strong temporal consistency.
  4. Sora-2 / Text-to-Video Pro — Pro-grade steerability and long-range coherence for complex sequences.
  5. GPT-Image-1 / Text-to-Image — Fast, prompt-faithful images with editability and tool-friendly outputs.
  6. DALL·E 3 — Clean composition and rich detail for concepting and illustration.
  7. DALL·E 2 — Lightweight text-to-image for quick drafts and style exploration.
  8. Sora (legacy) — Earlier Sora generation for baseline motion tests and rapid previews.
  9. Openai-whisper — High-accuracy multilingual speech recognition model for precise transcription with automatic language detection and punctuation.
  10. Openai-whisper-turbo — Optimized Whisper variant delivering the same accuracy with significantly faster transcription speed for real-time and large-scale use.
  11. Openai/gpt-image-1-mini/text-to-image generates high-quality images directly from text prompts with GPT-5-level understanding and efficiency, ideal for creative and design tasks.
  12. Openai/gpt-image-1-mini/edit enables intelligent image editing and refinement via natural-language instructions, preserving style and composition while applying precise changes.
  13. Openai/gpt-image-1-high-fidelity delivers ultra-detailed, photorealistic image generation powered by GPT-5, offering superior texture, lighting, and realism for professional-grade creative and design applications.
  14. Openai/gpt-image-1.5/text-to-image generates high-quality images from natural-language prompts with cost-efficient performance, producing coherent composition and clean aesthetics for UI concepts, marketing visuals, and fast creative ideation.
  15. Openai/gpt-image-1.5/text-to-image delivers high-quality text-to-image generation with strong prompt understanding and optimized synthesis, enabling rapid iteration and scalable visual production for design, prototyping, and creative workflows.
  16. Openai/gpt-image-2/edit enables high-fidelity image editing from natural-language instructions and reference images, preserving visual coherence, stylistic consistency, and fine-grained detail for marketing assets, design refinement, and fast creative iteration.
  17. Openai/gpt-image-2/text-to-image delivers high-quality text-to-image generation with strong prompt adherence, clean composition, and polished aesthetics, enabling scalable visual creation for UI concepts, campaign assets, and rapid creative ideation.

Why OpenAI Models?

  1. State-of-the-art quality — Physics-aware video, synchronized audio, and high-fidelity images with strong prompt adherence.
  2. End-to-end workflow — Text-to-image, image-to-video, and text-to-video in one stack; smooth handoff between models.
  3. Pro-grade control — Seeds, duration/aspect, camera language, and edit ops for consistent, repeatable results.
  4. Wide style range — From photoreal and documentary to anime, illustration, and cinematic looks—without plastic over-sharpening.

OpenAI Models API — Preise und Performance

Nutzen Sie jedes Modell der OpenAI Models-Sammlung über eine einzige REST-API. Bezahlen Sie pro Generierung — keine Abos, keine Mindestbeträge — mit branchenführender Latenz auf einer Infrastruktur mit 99,9 % Verfügbarkeit.

Warum OpenAI Models auf WaveSpeedAI ausführen

Transparente Preise

Abrechnung pro Aufruf für jedes OpenAI Models-Modell. Der Preis ist auf jeder Modellseite ausgewiesen — keine Plattformgebühren obendrauf.

Auf niedrige Latenz optimiert

Die meisten OpenAI Models-Bildmodelle laufen in unter 2 Sekunden. Video- und 3D-Modelle sind mehrfach schneller als selbst gehostete Alternativen.

99,9 % Verfügbarkeit

Multi-Region-Failover und automatische Wiederholungen halten Ihren Produktionsverkehr online — auch bei Anbieter-Ausfällen.

Häufig gestellte Fragen

Wie viel kostet die OpenAI Models-API?+

Jedes Modell hat seinen eigenen Preis pro Aufruf, der auf der Modellseite angegeben ist. Wir rechnen pro erfolgreicher Generierung ab — ohne Abogebühren oder Mindestbeträge.

Wie schnell sind OpenAI Models-Modelle auf WaveSpeedAI?+

Bildmodelle in dieser Sammlung sind typischerweise in unter 2 Sekunden fertig. Video- und 3D-Modelle hängen von Dauer und Auflösung ab, sind aber meist mehrfach schneller als selbst gehostete Läufe.

Kann ich die API ohne Kreditkarte testen?+

Ja — jedes Konto erhält bei der Anmeldung 20 $ Startguthaben, was Hunderte von Aufrufen über die meisten OpenAI Models-Modelle abdeckt.

Gibt es Rate-Limits?+

Standardkonten haben großzügige Limits für gleichzeitige Jobs. Enterprise-Pläne bieten individuelle RPM, höhere Parallelität und reservierte Kapazität — bei Interesse den Vertrieb kontaktieren.