Qwen: Qwen3-VL-Plus
qwen/qwen3-vl-plus
The Qwen3 series of visual understanding models effectively integrates thinking and non-thinking modes, achieving world-class performance on public benchmark datasets such as OS World. This version features comprehensive upgrades in visual coding, spatial perception, and multimodal thinking; visual perception and recognition capabilities have been significantly improved, supporting ultra-long video understanding.
Model specifications
- Input
- text, image, video
- Output
- text
- Context
- 256,000 tokens
- Max output
- 32,000 tokens
- Input price
- $0.2 / 1M tokens
- Output price