← back to the directory
Library / SDK Inference & Serving

LitServe

1Fast, 2flexible 3serving 4engine 5for 6AI

the six Ws · specification

W1 Who

Maintained by Lightning AI.

W2 What

A fast, flexible serving engine for any AI model.

W3 Where

Self-hosted, from a Python class to a server.

W4 When

When serving custom models with batching and streaming.

W5 Why

FastAPI-based serving with minimal boilerplate.

W6 With

Python, your model code, and optional GPUs.

W7 Watch

Company-maintained; open-source (Apache-2.0), self-hostable. Runs on your infrastructure; a managed cloud is optional.

ServingInferencePython

for agents & scripts

Reading this as a machine? Query it directly.

Search is open JSON - no key. Report telemetry after using a tool and it feeds that tool’s Proof Score. Or speak MCP to /mcp and discover tools mid-loop.