← back to the directory
Model Inference & Serving

CogVLM2

1Zhipu 2AI's 3visual 4expert 5language 6model

the six Ws · specification

W1 Who

Zhipu AI (THUDM) built CogVLM2 as a successor to the original CogVLM.

W2 What

A vision-language model using a visual expert module to fuse image and text features deeply.

W3 Where

Released on Hugging Face and GitHub under the THUDM organization.

W4 When

CogVLM2 was released in May 2024, with CogVLM originally launched in 2023.

W5 Why

It pioneered deep visual-language feature fusion, achieving strong results on OCR and chart-understanding benchmarks.

W6 With

Runs on PyTorch and transformers, built on a Llama3-based language backbone with a dedicated visual expert.

W7 Watch

Open-weight release by Zhipu AI/THUDM; free for research and, with registration, commercial use under its own license.

vision-languagemultimodalzhipu-aichinese-lab

for agents & scripts

Reading this as a machine? Query it directly.

Search is open JSON - no key. Report telemetry after using a tool and it feeds that tool’s Proof Score. Or speak MCP to /mcp and discover tools mid-loop.