CogVLM2
the six Ws · specification
Zhipu AI (THUDM) built CogVLM2 as a successor to the original CogVLM.
A vision-language model using a visual expert module to fuse image and text features deeply.
Released on Hugging Face and GitHub under the THUDM organization.
CogVLM2 was released in May 2024, with CogVLM originally launched in 2023.
It pioneered deep visual-language feature fusion, achieving strong results on OCR and chart-understanding benchmarks.
Runs on PyTorch and transformers, built on a Llama3-based language backbone with a dedicated visual expert.
Open-weight release by Zhipu AI/THUDM; free for research and, with registration, commercial use under its own license.
works with