← back to the directory
Library / SDK Observability & Evaluation

lighteval

1All-in-one 2toolkit 3for 4evaluating 5language 6models

the six Ws · specification

W1 Who

Hugging Face users and ML teams evaluating LLMs across multiple inference backends.

W2 What

A Python library for evaluating language models across vLLM, Transformers and Nanotron backends with hundreds of tasks.

W3 Where

Self-hosted library, runs locally, on Hugging Face Spaces, or in the cloud.

W4 When

Released by Hugging Face in 2024 and actively developed through 2026.

W5 Why

Consolidates fragmented evaluation tooling into one fast, extensible toolkit tied to the Hugging Face ecosystem.

W6 With

Requires Python and integrates with Hugging Face Transformers, Accelerate and vLLM.

W7 Watch

Open source under MIT, self-hosted, no credentials for local use, maintained by Hugging Face, over 2.5k GitHub stars.

llm-evaluationmulti-backendhugging-facebenchmarkingpython-library

for agents & scripts

Reading this as a machine? Query it directly.

Search is open JSON - no key. Report telemetry after using a tool and it feeds that tool’s Proof Score. Or speak MCP to /mcp and discover tools mid-loop.