qwen/qwen3.6-35b-a3b
262,144 context · $0.25/M input tokens · $2.00/M output tokens
Qwen3.6-35B-A3B is an open-weight multimodal model from Alibaba Cloud with 35B total parameters and 3B active parameters per token. It uses a hybrid sparse Mixture-of-Experts architecture that combines Gated DeltaNet linear attention with standard attention layers, giving it strong efficiency for coding, agentic workflows, long-context reasoning, and multimodal understanding. The model supports text, image, and video inputs, a 262K-token context window, thinking and non-thinking modes, function calling, and structured outputs.
사용량 기반 과금
선결제 없이 사용한 만큼만 지불
다음 코드 예시를 사용해 API와 연동하세요:
from openai import OpenAI
client = OpenAI(
api_key="YOUR_API_KEY",
base_url="https://llm.wavespeed.ai/v1"
)
response = client.chat.completions.create(
model="qwen/qwen3.6-35b-a3b",
messages=[
{"role": "user", "content": "Hello!"}
]
)
print(response.choices[0].message.content)Qwen3.6-35B-A3B is an open-weight multimodal model from Alibaba Cloud with 35B total parameters and 3B active parameters per token. It uses a hybrid sparse Mixture-of-Experts architecture combining Gated DeltaNet linear attention with standard attention layers, giving it strong efficiency for coding, agentic workflows, long-context reasoning, and multimodal understanding.
| Specification | Value |
|---|---|
| Provider | alibaba |
| Model Type | Chat Completions model |
| Architecture | Sparse MoE, 35B total / 3B active |
| Attention | Gated DeltaNet + standard attention |
| Modalities | text+image+video->text |
| Context Window | 262,144 tokens |
| Max Input | Not listed |
| Max Output | Not listed |
| Input | Text, Image, Video |
| Output | Text |
| Vision | Supported |
| Function Calling | Supported |
| Structured Outputs | Supported |
| Thinking Mode | Supported |
| Release | April 2026 |
| Token Type | Cost |
|---|---|
| Input | $0.149 per million tokens |
| Output | $1.00 per million tokens |
Base URL: https://llm.wavespeed.ai/v1
API Endpoint: chat/completions
Model ID: qwen/qwen3.6-35b-a3b
from openai import OpenAI
client = OpenAI(
api_key="YOUR_API_KEY",
base_url="https://llm.wavespeed.ai/v1"
)
response = client.chat.completions.create(
model="qwen/qwen3.6-35b-a3b",
messages=[{"role": "user", "content": "Hello!"}]
)
print(response.choices[0].message.content)
curl https://llm.wavespeed.ai/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer YOUR_API_KEY" \
-d '{
"model": "qwen/qwen3.6-35b-a3b",
"messages": [{"role": "user", "content": "Hello!"}]
}'
qwen/qwen3.6-35b-a3b
Qwen3.6-35B-A3B is an open-weight multimodal model from Alibaba Cloud with 35B total parameters and 3B active parameters per token. It uses a hybrid sparse Mixture-of-Experts architecture that combines Gated DeltaNet linear attention with standard attention layers, giving it strong efficiency for coding, agentic workflows, long-context reasoning, and multimodal understanding. The model supports text, image, and video inputs, a 262K-token context window, thinking and non-thinking modes, function calling, and structured outputs.
입력
$0.25 /M
출력
$2 /M
컨텍스트
262K
Vision
지원
도구 사용
지원
통합 API를 통해 Qwen3.6 35b A3b 액세스 — OpenAI 호환, 콜드 스타트 없음, 투명한 가격.
WaveSpeedAI 가격: 입력 토큰 100만 개당 $0.25, 출력 토큰 100만 개당 $2.00. 프롬프트 캐싱과 배치 처리는 별도로 청구되며 긴 반복 작업에서 실질 비용을 줄여 줍니다.
Qwen3.6 35b A3b은 요청당 최대 262K 컨텍스트 토큰과 최대 — 출력 토큰을 지원합니다.
네. WaveSpeedAI는 OpenAI 호환 엔드포인트 https://llm.wavespeed.ai/v1을 통해 Qwen3.6 35b A3b을 제공합니다. 공식 OpenAI SDK의 base URL을 이 주소로 변경하고 WaveSpeedAI API 키를 사용하면 코드 변경 없이 사용할 수 있습니다.
WaveSpeedAI에 로그인하고 Access Keys에서 API 키를 만든 다음, 위에 표시된 모델 ID로 https://llm.wavespeed.ai/v1/chat/completions에 요청을 보내세요. 신규 계정은 Qwen3.6 35b A3b을 평가할 수 있는 무료 크레딧을 받습니다.