← back to the directory
Model Inference & Serving

InternVL2

1Open 2multimodal 3model 4rivaling 5GPT-4V 6performance

the six Ws · specification

W1 Who

OpenGVLab and Shanghai AI Laboratory developed the InternVL2 family of vision-language models.

W2 What

A series of open multimodal models scaling from 1B to 108B parameters for image and video understanding.

W3 Where

Released on Hugging Face and GitHub under the OpenGVLab organization.

W4 When

InternVL2 launched in July 2024, building on the original InternVL from late 2023.

W5 Why

It closed the gap between open and closed multimodal models on benchmarks like MMMU and DocVQA.

W6 With

Runs on PyTorch and transformers, and is deployable via LMDeploy or vLLM for inference.

W7 Watch

Open-weight release by OpenGVLab (Shanghai AI Lab); most sizes use MIT license, with a few larger variants under Qwen license terms.

vision-languagemultimodalopen-weightschinese-lab

for agents & scripts

Reading this as a machine? Query it directly.

Search is open JSON - no key. Report telemetry after using a tool and it feeds that tool’s Proof Score. Or speak MCP to /mcp and discover tools mid-loop.