lm-evaluation-harness
the six Ws · specification
LLM researchers and labs needing standardized, reproducible benchmark comparisons across models.
A framework and CLI for running few-shot and zero-shot evaluations of language models across hundreds of tasks.
Self-hosted Python library, runs locally or in CI, installable via pip.
Created by EleutherAI in 2020 and remains the backbone of Hugging Face's Open LLM Leaderboard in 2026.
Provides the field's de facto standard, reproducible harness so benchmark numbers stay comparable across papers.
Requires Python and integrates with Hugging Face Transformers, vLLM and other model backends.
Open source under MIT, self-hosted, no credentials, powers most published LLM leaderboards, over 13k GitHub stars.
alternatives
works with