DeepEval
the six Ws · specification
Maintained by Confident AI.
A pytest-like framework offering LLM evaluation metrics for testing model and RAG outputs.
A Python library run locally or in CI, with an optional cloud dashboard.
Reach for it when writing assertions and regression tests over LLM outputs.
Brings software-style unit testing to LLM quality with ready-made metrics.
Works with pytest, LangChain, LlamaIndex, and any LLM provider.
Maintained by Confident AI; open-source (Apache-2.0) library run locally. Evaluation stays local unless you opt into the Confident AI cloud (open-core).
alternatives
works with