qwen/qwen3.6-flash
출시일: 2026-04-27
1,000,000 context · $0.25/M input tokens · $1.50/M output tokens
Qwen3.6 Flash is a fast, efficient multimodal language model from Alibaba’s Qwen 3.6 series. It supports text, image, and video inputs with a 1M-token context window and up to 64K output tokens. The model is designed for high-throughput chat, lightweight agent workflows, long-document understanding, visual reasoning, summarization, extraction, and cost-sensitive production workloads. It supports thinking mode, function calling, built-in tools, structured outputs, and batch calling.
사용량 기반 과금
선결제 없이 사용한 만큼만 지불
다음 코드 예시를 사용해 API와 연동하세요:
import OpenAI from 'openai';
if (!process.env.WAVESPEED_API_KEY) throw new Error('Set WAVESPEED_API_KEY');
const client = new OpenAI({
apiKey: process.env.WAVESPEED_API_KEY,
baseURL: 'https://llm.wavespeed.ai/v1',
timeout: 120_000,
maxRetries: 2,
});
try {
const response = await client.chat.completions.create({
model: 'qwen/qwen3.6-flash',
messages: [{ role: 'user', content: 'Hello!' }],
});
console.log(response.choices[0]?.message?.content ?? '');
} catch (error) {
console.error('LLM request failed:', error);
process.exitCode = 1;
}import OpenAI from 'openai';
if (!process.env.WAVESPEED_API_KEY) throw new Error('Set WAVESPEED_API_KEY');
const client = new OpenAI({
apiKey: process.env.WAVESPEED_API_KEY,
baseURL: 'https://llm.wavespeed.ai/v1',
timeout: 120_000,
maxRetries: 2,
});
try {
const response = await client.chat.completions.create({
model: 'qwen/qwen3.6-flash',
messages: [{ role: 'user', content: 'Hello!' }],
});
console.log(response.choices[0]?.message?.content ?? '');
} catch (error) {
console.error('LLM request failed:', error);
process.exitCode = 1;
}Qwen3.6 Flash is a fast, efficient multimodal language model from Alibaba’s Qwen 3.6 series. It supports text, image, and video inputs with a 1M-token context window, making it a strong fit for high-volume chat, lightweight agents, long-document workflows, visual understanding, summarization, and structured extraction.
| Specification | Value |
|---|---|
| Provider | alibaba |
| Model Type | Chat Completions model |
| Architecture | text+image+video->text |
| Context Window | 1,000,000 tokens |
| Max Input | 934,464 tokens |
| Max Output | 65,536 tokens |
| Thinking Budget | 128K tokens |
| Input | Text, Image, Video |
| Output | Text |
| Vision | Supported |
| Function Calling | Supported |
| Built-in Tools | Supported |
| Structured Outputs | Supported |
| Batch Calling | Supported |
Base URL: https://llm.wavespeed.ai/v1 API Endpoint: chat/completions Model ID: qwen/qwen3.6-flash
from openai import OpenAI
client = OpenAI(
api_key="YOUR_API_KEY",
base_url="https://llm.wavespeed.ai/v1"
)
response = client.chat.completions.create(
model="qwen/qwen3.6-flash",
messages=[{"role": "user", "content": "Hello!"}]
)
print(response.choices[0].message.content)qwen/qwen3.6-flash
Qwen3.6 Flash is a fast, efficient multimodal language model from Alibaba’s Qwen 3.6 series. It supports text, image, and video inputs with a 1M-token context window and up to 64K output tokens. The model is designed for high-throughput chat, lightweight agent workflows, long-document understanding, visual reasoning, summarization, extraction, and cost-sensitive production workloads. It supports thinking mode, function calling, built-in tools, structured outputs, and batch calling.
입력
$0.25 /M
출력
$1.5 /M
컨텍스트
1000K
최대 출력
66K
Vision
지원
도구 사용
지원
통합 API를 통해 Qwen3.6 Flash 액세스 — OpenAI 호환, 콜드 스타트 없음, 투명한 가격.
WaveSpeedAI 가격: 입력 토큰 100만 개당 $0.25, 출력 토큰 100만 개당 $1.50. 프롬프트 캐싱과 배치 처리는 별도로 청구되며 긴 반복 작업에서 실질 비용을 줄여 줍니다.
Qwen3.6 Flash은 요청당 최대 1000K 컨텍스트 토큰과 최대 66K 출력 토큰을 지원합니다.
WaveSpeedAI는 https://llm.wavespeed.ai/v1의 OpenAI 호환 Chat Completions 인터페이스를 통해 Qwen3.6 Flash을 제공합니다. 대부분의 OpenAI SDK 클라이언트는 base URL과 API 키를 변경해 사용할 수 있으며, 선택 필드는 모델에 따라 다릅니다.
WaveSpeedAI에 로그인하고 Access Keys에서 API 키를 만든 다음, 위에 표시된 모델 ID로 https://llm.wavespeed.ai/v1/chat/completions에 요청을 보내세요. 제공 여부, 기능 및 가격은 최신 모델 카탈로그를 확인하세요.