Library / SDK
Inference & Serving
Triton Inference Server
1NVIDIA 2server 3for 4high-performance 5model 6inference
the six Ws · specification
W1
Who
Maintained by NVIDIA.
W2
What
Serves models from many frameworks with high performance.
W3
Where
Self-hosted on GPUs or CPUs, often in Kubernetes.
W4
When
When serving mixed model types in production.
W5
Why
Standardizes serving with batching, ensembles, and metrics.
W6
With
GPUs or CPUs, model repositories, and Docker.
W7
Watch
Vendor-maintained (NVIDIA); open-source (BSD-3-Clause), self-hostable. Runs on your own hardware; nothing leaves your infrastructure.
alternatives