Library / SDK
Inference & Serving
LitServe
1Fast, 2flexible 3serving 4engine 5for 6AI
the six Ws · specification
W1
Who
Maintained by Lightning AI.
W2
What
A fast, flexible serving engine for any AI model.
W3
Where
Self-hosted, from a Python class to a server.
W4
When
When serving custom models with batching and streaming.
W5
Why
FastAPI-based serving with minimal boilerplate.
W6
With
Python, your model code, and optional GPUs.
W7
Watch
Company-maintained; open-source (Apache-2.0), self-hostable. Runs on your infrastructure; a managed cloud is optional.
alternatives