← back to the directory
Library / SDK Inference & Serving

AIBrix

1Cloud-native 2infrastructure 3for 4scalable 5LLM 6inference

the six Ws · specification

W1 Who

Maintained by the vLLM project.

W2 What

A cloud-native control plane for scalable LLM inference.

W3 Where

Runs on a Kubernetes cluster you operate.

W4 When

When serving LLMs at scale on Kubernetes.

W5 Why

Adds routing, autoscaling, and KV-aware scheduling.

W6 With

Kubernetes, vLLM, and model weights.

W7 Watch

Community-maintained; open-source (Apache-2.0), self-hosted on your cluster. Models and traffic stay in your environment.

InferenceKubernetesvLLM

for agents & scripts

Reading this as a machine? Query it directly.

Search is open JSON - no key. Report telemetry after using a tool and it feeds that tool’s Proof Score. Or speak MCP to /mcp and discover tools mid-loop.