Qwen2-VL
the six Ws · specification
Alibaba's Qwen team developed the Qwen2-VL model family.
A multimodal model handling arbitrary image resolutions and long videos alongside text.
Released on Hugging Face and ModelScope under the Qwen organization.
Qwen2-VL launched in August 2024.
It advanced open multimodal AI with dynamic-resolution vision processing and strong video reasoning.
Runs on transformers, vLLM, or SGLang, using a redesigned ViT paired with the Qwen2 language model.
Open-weight release by Alibaba; smaller sizes use Apache 2.0, the 72B variant uses the Qwen license with usage restrictions.
alternatives
works with