google/gemini-3.5-flash
Yayın tarihi: 2026-05-19
1,048,576 context · $1.50/M input tokens · $9.00/M output tokens
Gemini 3.5 Flash is Google’s high-efficiency multimodal model, delivering near-Pro-level reasoning and coding capabilities with Flash-class speed and cost efficiency. It is purpose-built for advanced coding workflows and parallel agentic execution, while supporting a wide range of input modalities including text, images, video, audio, and PDFs.
The model defaults to a medium reasoning mode to balance latency, quality, and cost, while also offering configurable thinking levels — minimal, low, medium, and high — for more precise performance and efficiency tuning across different workloads.
Kullandıkça öde
Ön ödeme yok, yalnızca kullandığınız kadar ödeyin
API'mizle entegre etmek için aşağıdaki kod örneklerini kullanın:
import OpenAI from 'openai';
if (!process.env.WAVESPEED_API_KEY) throw new Error('Set WAVESPEED_API_KEY');
const client = new OpenAI({
apiKey: process.env.WAVESPEED_API_KEY,
baseURL: 'https://llm.wavespeed.ai/v1',
timeout: 120_000,
maxRetries: 2,
});
try {
const response = await client.chat.completions.create({
model: 'google/gemini-3.5-flash',
messages: [{ role: 'user', content: 'Hello!' }],
});
console.log(response.choices[0]?.message?.content ?? '');
} catch (error) {
console.error('LLM request failed:', error);
process.exitCode = 1;
}import OpenAI from 'openai';
if (!process.env.WAVESPEED_API_KEY) throw new Error('Set WAVESPEED_API_KEY');
const client = new OpenAI({
apiKey: process.env.WAVESPEED_API_KEY,
baseURL: 'https://llm.wavespeed.ai/v1',
timeout: 120_000,
maxRetries: 2,
});
try {
const response = await client.chat.completions.create({
model: 'google/gemini-3.5-flash',
messages: [{ role: 'user', content: 'Hello!' }],
});
console.log(response.choices[0]?.message?.content ?? '');
} catch (error) {
console.error('LLM request failed:', error);
process.exitCode = 1;
}Gemini 3.5 Flash is Google’s high-efficiency multimodal model, delivering near-Pro-level reasoning and coding capabilities with Flash-class speed and cost efficiency. It is purpose-built for advanced coding workflows and parallel agentic execution, while supporting a wide range of input modalities including text, images, video, audio, and PDFs.
The model defaults to a medium reasoning mode to balance latency, quality, and cost, while also offering configurable thinking levels — minimal, low, medium, and high — for more precise performance and efficiency tuning across different workloads.
| Specification | Value |
|---|---|
| Provider | |
| Model Type | Chat Completions model |
| Architecture | text+image+file+audio+video->text |
| Context Window | 1048576 tokens |
| Max Input | 983040 tokens |
| Max Output | 65536 tokens |
| Input | Text, Image, Video, file, Audio |
| Output | Text |
| Vision | Supported |
| Function Calling | Supported |
| Structured Outputs | Supported |
Base URL: https://llm.wavespeed.ai/v1 API Endpoint: chat/completions Model ID: google/gemini-3.5-flash
from openai import OpenAI
client = OpenAI(
api_key="YOUR_API_KEY",
base_url="https://llm.wavespeed.ai/v1"
)
response = client.chat.completions.create(
model="google/gemini-3.5-flash",
messages=[{"role": "user", "content": "Hello!"}]
)
print(response.choices[0].message.content)
curl https://llm.wavespeed.ai/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer YOUR_API_KEY" \
-d '{
"model": "google/gemini-3.5-flash",
"messages": [{"role": "user", "content": "Hello!"}]
}'google/gemini-3.5-flash
Gemini 3.5 Flash is Google’s high-efficiency multimodal model, delivering near-Pro-level reasoning and coding capabilities with Flash-class speed and cost efficiency. It is purpose-built for advanced coding workflows and parallel agentic execution, while supporting a wide range of input modalities including text, images, video, audio, and PDFs. The model defaults to a medium reasoning mode to balance latency, quality, and cost, while also offering configurable thinking levels — minimal, low, medium, and high — for more precise performance and efficiency tuning across different workloads.
Giriş
$1.5 /M
Çıkış
$9 /M
Bağlam
1049K
Maks. Çıkış
66K
Vision
Destekleniyor
Araç Kullanımı
Destekleniyor
Birleşik API'miz aracılığıyla Gemini 3.5 Flash'e erişin — OpenAI uyumlu, soğuk başlatma yok, şeffaf fiyatlandırma.
WaveSpeedAI fiyatlandırması: milyon giriş tokenı başına $1.50 ve milyon çıkış tokenı başına $9.00. Prompt caching ve toplu işleme ayrı faturalanır ve uzun, tekrar eden yüklerde etkin maliyeti düşürür.
Gemini 3.5 Flash istek başına 1049K bağlam tokenını ve 66K çıkış tokenını destekler.
WaveSpeedAI, Gemini 3.5 Flash modelini https://llm.wavespeed.ai/v1 adresindeki OpenAI uyumlu Chat Completions arayüzü üzerinden sunar. Çoğu OpenAI SDK istemcisinde base URL ve API anahtarını değiştirmek yeterlidir; isteğe bağlı alanlar modele bağlıdır.
WaveSpeedAI’a giriş yapın, Access Keys bölümünde bir API anahtarı oluşturun ve yukarıdaki model id ile https://llm.wavespeed.ai/v1/chat/completions adresine istek gönderin. Güncel kullanılabilirlik, özellikler ve fiyatlar için model kataloğunu kontrol edin.