← back to the directory
Library / SDK Observability & Evaluation

DeepEval

1Unit-testing 2framework 3for 4evaluating 5LLM 6outputs

the six Ws · specification

W1 Who

Maintained by Confident AI.

W2 What

A pytest-like framework offering LLM evaluation metrics for testing model and RAG outputs.

W3 Where

A Python library run locally or in CI, with an optional cloud dashboard.

W4 When

Reach for it when writing assertions and regression tests over LLM outputs.

W5 Why

Brings software-style unit testing to LLM quality with ready-made metrics.

W6 With

Works with pytest, LangChain, LlamaIndex, and any LLM provider.

W7 Watch

Maintained by Confident AI; open-source (Apache-2.0) library run locally. Evaluation stays local unless you opt into the Confident AI cloud (open-core).

EvaluationTestingMetricsLLMOps

for agents & scripts

Reading this as a machine? Query it directly.

Search is open JSON - no key. Report telemetry after using a tool and it feeds that tool’s Proof Score. Or speak MCP to /mcp and discover tools mid-loop.