← back to the directory
Library / SDK Observability & Evaluation

OpenCompass

1LLM 2evaluation 3platform 4across 5many 6datasets

the six Ws · specification

W1 Who

Researchers benchmarking LLMs such as GPT-4, Llama, Qwen and Claude across broad task suites.

W2 What

An open evaluation platform supporting over 100 datasets and many model families, with a public leaderboard.

W3 Where

Self-hosted Python toolkit, with results also published to the public OpenCompass leaderboard site.

W4 When

Launched by Shanghai AI Lab in 2023 and continuously updated with new model support through 2026.

W5 Why

Offers one of the broadest, most systematic multi-dataset evaluation suites for comparing LLMs fairly.

W6 With

Requires Python and connects to model APIs or local Transformers/vLLM backends for inference.

W7 Watch

Open source under Apache-2.0, self-hosted, API keys needed only when evaluating closed commercial models, over 7k GitHub stars.

llm-evaluationleaderboardmulti-datasetbenchmarkingreproducibility

for agents & scripts

Reading this as a machine? Query it directly.

Search is open JSON - no key. Report telemetry after using a tool and it feeds that tool’s Proof Score. Or speak MCP to /mcp and discover tools mid-loop.