Seedream 5.0 Pro yayında | Görsel Üretici'de deneyin →
google
google/gemini-3.5-flash

google/gemini-3.5-flash

Yayın tarihi: 2026-05-19

1,048,576 context · $1.50/M input tokens · $9.00/M output tokens

Gemini 3.5 Flash is Google’s high-efficiency multimodal model, delivering near-Pro-level reasoning and coding capabilities with Flash-class speed and cost efficiency. It is purpose-built for advanced coding workflows and parallel agentic execution, while supporting a wide range of input modalities including text, images, video, audio, and PDFs.

The model defaults to a medium reasoning mode to balance latency, quality, and cost, while also offering configurable thinking levels — minimal, low, medium, and high — for more precise performance and efficiency tuning across different workloads.

Fiyatlandırma

Kullandıkça öde

Ön ödeme yok, yalnızca kullandığınız kadar ödeyin

Giriş$1.50 / M Tokens
Çıkış$9.00 / M Tokens
Cache Read$0.15 / M Tokens
Cache Write$0.08 / M Tokens

Modeli dene

google/gemini-3.5-flash
Çevrimiçi
google
Merhaba! Yardımcı bir yapay zeka asistanıyım. Size nasıl yardımcı olabilirim?
Bu modeli yerel bir coding agent içinde kullanmaya hazır mısınız?Agent setup

API Kullanımı

API'mizle entegre etmek için aşağıdaki kod örneklerini kullanın:

import OpenAI from 'openai';

if (!process.env.WAVESPEED_API_KEY) throw new Error('Set WAVESPEED_API_KEY');
const client = new OpenAI({
  apiKey: process.env.WAVESPEED_API_KEY,
  baseURL: 'https://llm.wavespeed.ai/v1',
  timeout: 120_000,
  maxRetries: 2,
});

try {
  const response = await client.chat.completions.create({
    model: 'google/gemini-3.5-flash',
    messages: [{ role: 'user', content: 'Hello!' }],
  });
  console.log(response.choices[0]?.message?.content ?? '');
} catch (error) {
  console.error('LLM request failed:', error);
  process.exitCode = 1;
}

Model Tanıtımı

Google: Gemini 3.5 Flash

Gemini 3.5 Flash is Google’s high-efficiency multimodal model, delivering near-Pro-level reasoning and coding capabilities with Flash-class speed and cost efficiency. It is purpose-built for advanced coding workflows and parallel agentic execution, while supporting a wide range of input modalities including text, images, video, audio, and PDFs.

The model defaults to a medium reasoning mode to balance latency, quality, and cost, while also offering configurable thinking levels — minimal, low, medium, and high — for more precise performance and efficiency tuning across different workloads.


Why It Looks Great

  • text+image+file+audio+video->text architecture for Text, Image, Video, file, Audio to Text workloads
  • 1048576 context window for long prompts, document analysis, and multi-turn workflows
  • Competitive pricing at $1.5/$9 per million tokens
  • Vision input support for image understanding and multimodal tasks
  • Function calling and tool-use support for agentic application workflows
  • Structured output support for JSON responses and schema-constrained generation

Key Features

  • Context Window: 1048576 tokens
  • Max Input: 983040 tokens
  • Max Output: 65536 tokens
  • Vision: Supported
  • Function Calling: Supported
  • Structured Outputs: Supported
  • Image Generation: Not listed
  • Audio Input: Supported
  • Supported Parameters: include_reasoning, max_tokens, reasoning, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_p

Specifications

SpecificationValue
Providergoogle
Model TypeChat Completions model
Architecturetext+image+file+audio+video->text
Context Window1048576 tokens
Max Input983040 tokens
Max Output65536 tokens
InputText, Image, Video, file, Audio
OutputText
VisionSupported
Function CallingSupported
Structured OutputsSupported

How to Use

  1. Write your prompt - describe the task, provide context, and specify the desired output format.
  2. Submit - the model processes your request and returns the response.

API Integration

Base URL: https://llm.wavespeed.ai/v1 API Endpoint: chat/completions Model ID: google/gemini-3.5-flash


API Usage

Python SDK

from openai import OpenAI

client = OpenAI(
    api_key="YOUR_API_KEY",
    base_url="https://llm.wavespeed.ai/v1"
)

response = client.chat.completions.create(
    model="google/gemini-3.5-flash",
    messages=[{"role": "user", "content": "Hello!"}]
)

print(response.choices[0].message.content)

cURL

curl https://llm.wavespeed.ai/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -d '{
    "model": "google/gemini-3.5-flash",
    "messages": [{"role": "user", "content": "Hello!"}]
  }'

Bilgi

Sağlayıcıgoogle
Türllm

Desteklenen İşlevsellik

Giriş
MetinGörselSes
Çıkış
Metin
Bağlam1,048,576
Maks. Çıkış65,536
Vision✓ Destekleniyor
Function Calling✓ Destekleniyor

API Erişim Kılavuzu

Base URLhttps://llm.wavespeed.ai/v1
API Endpointchat/completions
Model IDgoogle/gemini-3.5-flash

Gemini 3.5 Flash API

google/gemini-3.5-flash

Gemini 3.5 Flash is Google’s high-efficiency multimodal model, delivering near-Pro-level reasoning and coding capabilities with Flash-class speed and cost efficiency. It is purpose-built for advanced coding workflows and parallel agentic execution, while supporting a wide range of input modalities including text, images, video, audio, and PDFs. The model defaults to a medium reasoning mode to balance latency, quality, and cost, while also offering configurable thinking levels — minimal, low, medium, and high — for more precise performance and efficiency tuning across different workloads.

Giriş

$1.5 /M

Çıkış

$9 /M

Bağlam

1049K

Maks. Çıkış

66K

Vision

Destekleniyor

Araç Kullanımı

Destekleniyor

Gemini 3.5 Flash'i WaveSpeedAI'da deneyin

Birleşik API'miz aracılığıyla Gemini 3.5 Flash'e erişin — OpenAI uyumlu, soğuk başlatma yok, şeffaf fiyatlandırma.

Gemini 3.5 Flash hakkında sık sorulan sorular

Gemini 3.5 Flash API ücreti ne kadar?+

WaveSpeedAI fiyatlandırması: milyon giriş tokenı başına $1.50 ve milyon çıkış tokenı başına $9.00. Prompt caching ve toplu işleme ayrı faturalanır ve uzun, tekrar eden yüklerde etkin maliyeti düşürür.

Gemini 3.5 Flash'in bağlam penceresi nedir?+

Gemini 3.5 Flash istek başına 1049K bağlam tokenını ve 66K çıkış tokenını destekler.

Gemini 3.5 Flash OpenAI uyumlu mu?+

WaveSpeedAI, Gemini 3.5 Flash modelini https://llm.wavespeed.ai/v1 adresindeki OpenAI uyumlu Chat Completions arayüzü üzerinden sunar. Çoğu OpenAI SDK istemcisinde base URL ve API anahtarını değiştirmek yeterlidir; isteğe bağlı alanlar modele bağlıdır.

Gemini 3.5 Flash'e nasıl başlarım?+

WaveSpeedAI’a giriş yapın, Access Keys bölümünde bir API anahtarı oluşturun ve yukarıdaki model id ile https://llm.wavespeed.ai/v1/chat/completions adresine istek gönderin. Güncel kullanılabilirlik, özellikler ve fiyatlar için model kataloğunu kontrol edin.

İlgili LLM API'leri

Gemini 3.5 Flash | Google Efficient LLM API | WaveSpeedAI