HELM
the six Ws · specification
Academic researchers and labs wanting transparent, multi-metric evaluation of foundation models.
A Python framework from Stanford CRFM for holistic, reproducible evaluation of language and multimodal models across many scenarios and metrics.
Self-hosted Python package, with results published to the public crfm.stanford.edu HELM leaderboard.
Introduced by Stanford CRFM in 2022 and continuously expanded with new benchmarks since.
Pushes beyond single accuracy scores to measure fairness, robustness, calibration and efficiency together.
Requires Python and API access or local weights for the models being evaluated.
Open source under Apache-2.0, self-hosted, API keys needed for closed commercial models, backed by Stanford CRFM, over 2.8k GitHub stars.
alternatives