← back to the directory
Model Inference & Serving

Qwen2-VL

1Alibaba's 2vision-language 3model 4handles 5images, 6video

the six Ws · specification

W1 Who

Alibaba's Qwen team developed the Qwen2-VL model family.

W2 What

A multimodal model handling arbitrary image resolutions and long videos alongside text.

W3 Where

Released on Hugging Face and ModelScope under the Qwen organization.

W4 When

Qwen2-VL launched in August 2024.

W5 Why

It advanced open multimodal AI with dynamic-resolution vision processing and strong video reasoning.

W6 With

Runs on transformers, vLLM, or SGLang, using a redesigned ViT paired with the Qwen2 language model.

W7 Watch

Open-weight release by Alibaba; smaller sizes use Apache 2.0, the 72B variant uses the Qwen license with usage restrictions.

vision-languagemultimodalalibabavideo-understanding

for agents & scripts

Reading this as a machine? Query it directly.

Search is open JSON - no key. Report telemetry after using a tool and it feeds that tool’s Proof Score. Or speak MCP to /mcp and discover tools mid-loop.