← back to the directory
Model Inference & Serving

CogVideoX

1Open 2diffusion 3transformer 4generates 5coherent 6video

the six Ws · specification

W1 Who

Zhipu AI and Tsinghua University's KEG lab released CogVideoX.

W2 What

CogVideoX is an open-weight text-to-video diffusion transformer model designed for generating longer, more temporally coherent video clips.

W3 Where

Hosted on GitHub at THUDM/CogVideo and on Hugging Face under THUDM and zai-org.

W4 When

Released in August 2024, with the CogVideoX-5B variant following shortly after.

W5 Why

It was one of the first strong openly licensed text-to-video models rivaling early closed commercial offerings.

W6 With

Built on PyTorch and integrates with the Hugging Face diffusers library.

W7 Watch

The smaller CogVideoX-2B is released under the Apache 2.0 license while the 5B variant uses a custom commercial-restricted license; released by Zhipu AI.

text-to-videodiffusion transformerZhipu AIopen weights

for agents & scripts

Reading this as a machine? Query it directly.

Search is open JSON - no key. Report telemetry after using a tool and it feeds that tool’s Proof Score. Or speak MCP to /mcp and discover tools mid-loop.