リリース2026/02
生成速度144 トークン/秒
初動レイテンシ (TTFT)1.02s
入力料金$0.4/1M
出力料金$3.2/1M
ランキング
レビュー
Qwen3.5 is a multimodal-specialized model, ideal when images or video are central. Avoid using it for coding, general chat, or autonomous agents—opt for specialized models in those areas instead.
長所
- Outstanding multimodal processing—ranked #4 globally, with strong comprehension of integrated images, video, and text
- Stable logical reasoning—effective for complex data analysis and deep contextual understanding
- Consistent output quality—well-balanced between accuracy and speed
短所
- Weak in coding—ranked #42 with a low AIM score of 61.5, insufficient for code-heavy workloads
- Weak general conversation (ranked #67)—not a good choice for chatbots or standard Q&A
- Limited agentic capabilities—ranked #88 in AI Agent, not well-suited for autonomous multi-step tasks
ユースケース
Image and video processing, OCR, and visual content descriptionComplex document analysis featuring charts/images alongside textReasoning tasks requiring cross-modal contextual understanding
ガイド・動画
Visit qwen35.com to use it directly in the browser without setup. For local deployment, run it via SGLang or vLLM on H100/A100 GPUs (requires 8 GPUs). The model handles complex reasoning, multimodal processing (text, images, video), programming, and multi-step agent tasks. Tip: use thinking mode for challenging logic problems and enable tool calling for tasks requiring function calls.