← back to the directory
Library / SDK Observability & Evaluation

Inspect AI

1A 2framework 3for 4large 5language 6evals

the six Ws · specification

W1 Who

Maintained by the UK AI Safety Institute.

W2 What

A framework for building and running language model evals.

W3 Where

A Python library run locally or in CI.

W4 When

When rigorously evaluating model or agent capabilities.

W5 Why

Standardizes datasets, solvers, and scorers for evals.

W6 With

Python and the model providers under test.

W7 Watch

Government-maintained (UK AISI); open-source (MIT), self-runnable. Evals run locally and call whichever model providers you point them at.

EvaluationBenchmarksPython

for agents & scripts

Reading this as a machine? Query it directly.

Search is open JSON - no key. Report telemetry after using a tool and it feeds that tool’s Proof Score. Or speak MCP to /mcp and discover tools mid-loop.