Library / SDK
Inference & Serving
Xinference
1Serve 2any 3open 4model, 5one 6command
the six Ws · specification
W1
Who
Maintained by Xorbits (xorbitsai).
W2
What
Serves LLMs, embeddings, and multimodal models via one API.
W3
Where
Self-hosted on CPU or GPU clusters.
W4
When
When serving many model types behind one endpoint.
W5
Why
Unifies model serving with an OpenAI-compatible interface.
W6
With
Python, model weights, and optional GPUs.
W7
Watch
Company-maintained; open-source (Apache-2.0), self-hostable. Runs on your own hardware; models and traffic stay local.