InternVL2
the six Ws · specification
OpenGVLab and Shanghai AI Laboratory developed the InternVL2 family of vision-language models.
A series of open multimodal models scaling from 1B to 108B parameters for image and video understanding.
Released on Hugging Face and GitHub under the OpenGVLab organization.
InternVL2 launched in July 2024, building on the original InternVL from late 2023.
It closed the gap between open and closed multimodal models on benchmarks like MMMU and DocVQA.
Runs on PyTorch and transformers, and is deployable via LMDeploy or vLLM for inference.
Open-weight release by OpenGVLab (Shanghai AI Lab); most sizes use MIT license, with a few larger variants under Qwen license terms.