Library / SDK
Inference & Serving
TensorRT-LLM
1NVIDIA's 2optimized 3LLM 4inference 5on 6GPUs
the six Ws · specification
W1
Who
Maintained by NVIDIA.
W2
What
Optimizes and serves LLMs for maximum GPU performance.
W3
Where
Self-hosted on NVIDIA GPUs.
W4
When
When you need peak inference throughput on NVIDIA.
W5
Why
Compiles models to squeeze the most from NVIDIA hardware.
W6
With
NVIDIA GPUs, CUDA, and model weights.
W7
Watch
Vendor-maintained (NVIDIA); open-source (Apache-2.0), self-hostable. Runs on your GPUs; models and traffic stay local.
works with