← back to the directory
Library / SDK Observability & Evaluation

lm-evaluation-harness

1Standardized 2few-shot 3benchmarking 4framework 5for 6LLMs

the six Ws · specification

W1 Who

LLM researchers and labs needing standardized, reproducible benchmark comparisons across models.

W2 What

A framework and CLI for running few-shot and zero-shot evaluations of language models across hundreds of tasks.

W3 Where

Self-hosted Python library, runs locally or in CI, installable via pip.

W4 When

Created by EleutherAI in 2020 and remains the backbone of Hugging Face's Open LLM Leaderboard in 2026.

W5 Why

Provides the field's de facto standard, reproducible harness so benchmark numbers stay comparable across papers.

W6 With

Requires Python and integrates with Hugging Face Transformers, vLLM and other model backends.

W7 Watch

Open source under MIT, self-hosted, no credentials, powers most published LLM leaderboards, over 13k GitHub stars.

llm-benchmarkingfew-shot-evalstandardized-taskscli-toolopen-source

for agents & scripts

Reading this as a machine? Query it directly.

Search is open JSON - no key. Report telemetry after using a tool and it feeds that tool’s Proof Score. Or speak MCP to /mcp and discover tools mid-loop.