OpenCompass
the six Ws · specification
Researchers benchmarking LLMs such as GPT-4, Llama, Qwen and Claude across broad task suites.
An open evaluation platform supporting over 100 datasets and many model families, with a public leaderboard.
Self-hosted Python toolkit, with results also published to the public OpenCompass leaderboard site.
Launched by Shanghai AI Lab in 2023 and continuously updated with new model support through 2026.
Offers one of the broadest, most systematic multi-dataset evaluation suites for comparing LLMs fairly.
Requires Python and connects to model APIs or local Transformers/vLLM backends for inference.
Open source under Apache-2.0, self-hosted, API keys needed only when evaluating closed commercial models, over 7k GitHub stars.
alternatives