CogVideoX
the six Ws · specification
Zhipu AI and Tsinghua University's KEG lab released CogVideoX.
CogVideoX is an open-weight text-to-video diffusion transformer model designed for generating longer, more temporally coherent video clips.
Hosted on GitHub at THUDM/CogVideo and on Hugging Face under THUDM and zai-org.
Released in August 2024, with the CogVideoX-5B variant following shortly after.
It was one of the first strong openly licensed text-to-video models rivaling early closed commercial offerings.
Built on PyTorch and integrates with the Hugging Face diffusers library.
The smaller CogVideoX-2B is released under the Apache 2.0 license while the 5B variant uses a custom commercial-restricted license; released by Zhipu AI.
works with