← back to the directory
Library / SDK Observability & Evaluation

HELM

1Holistic 2reproducible 3framework 4for 5evaluating 6models

the six Ws · specification

W1 Who

Academic researchers and labs wanting transparent, multi-metric evaluation of foundation models.

W2 What

A Python framework from Stanford CRFM for holistic, reproducible evaluation of language and multimodal models across many scenarios and metrics.

W3 Where

Self-hosted Python package, with results published to the public crfm.stanford.edu HELM leaderboard.

W4 When

Introduced by Stanford CRFM in 2022 and continuously expanded with new benchmarks since.

W5 Why

Pushes beyond single accuracy scores to measure fairness, robustness, calibration and efficiency together.

W6 With

Requires Python and API access or local weights for the models being evaluated.

W7 Watch

Open source under Apache-2.0, self-hosted, API keys needed for closed commercial models, backed by Stanford CRFM, over 2.8k GitHub stars.

holistic-evaluationbenchmarkingtransparencyresearch-frameworkleaderboard

for agents & scripts

Reading this as a machine? Query it directly.

Search is open JSON - no key. Report telemetry after using a tool and it feeds that tool’s Proof Score. Or speak MCP to /mcp and discover tools mid-loop.