← back to the directory
Library / SDK Inference & Serving

Triton Inference Server

1NVIDIA 2server 3for 4high-performance 5model 6inference

the six Ws · specification

W1 Who

Maintained by NVIDIA.

W2 What

Serves models from many frameworks with high performance.

W3 Where

Self-hosted on GPUs or CPUs, often in Kubernetes.

W4 When

When serving mixed model types in production.

W5 Why

Standardizes serving with batching, ensembles, and metrics.

W6 With

GPUs or CPUs, model repositories, and Docker.

W7 Watch

Vendor-maintained (NVIDIA); open-source (BSD-3-Clause), self-hostable. Runs on your own hardware; nothing leaves your infrastructure.

InferenceServingNVIDIA

for agents & scripts

Reading this as a machine? Query it directly.

Search is open JSON - no key. Report telemetry after using a tool and it feeds that tool’s Proof Score. Or speak MCP to /mcp and discover tools mid-loop.