← back to the directory
Library / SDK Inference & Serving

Xinference

1Serve 2any 3open 4model, 5one 6command

the six Ws · specification

W1 Who

Maintained by Xorbits (xorbitsai).

W2 What

Serves LLMs, embeddings, and multimodal models via one API.

W3 Where

Self-hosted on CPU or GPU clusters.

W4 When

When serving many model types behind one endpoint.

W5 Why

Unifies model serving with an OpenAI-compatible interface.

W6 With

Python, model weights, and optional GPUs.

W7 Watch

Company-maintained; open-source (Apache-2.0), self-hostable. Runs on your own hardware; models and traffic stay local.

InferenceServingModels

for agents & scripts

Reading this as a machine? Query it directly.

Search is open JSON - no key. Report telemetry after using a tool and it feeds that tool’s Proof Score. Or speak MCP to /mcp and discover tools mid-loop.