← back to the directory
Library / SDK Inference & Serving

KubeAI

1Serve 2AI 3models 4on 5Kubernetes 6easily

the six Ws · specification

W1 Who

Maintained by Substratus.

W2 What

Serves LLMs and embeddings natively on Kubernetes.

W3 Where

Runs on a Kubernetes cluster you operate.

W4 When

When you need OpenAI-compatible inference on Kubernetes.

W5 Why

Autoscaling model serving with no external dependencies.

W6 With

Kubernetes and model weights.

W7 Watch

Company-maintained; open-source (Apache-2.0), self-hosted on your cluster. Models and traffic stay in your environment.

KubernetesServingInference

for agents & scripts

Reading this as a machine? Query it directly.

Search is open JSON - no key. Report telemetry after using a tool and it feeds that tool’s Proof Score. Or speak MCP to /mcp and discover tools mid-loop.