Library / SDK
Inference & Serving
LMDeploy
1Compress, 2deploy, 3and 4serve 5LLMs 6efficiently
the six Ws · specification
W1
Who
Maintained by the InternLM team (Shanghai AI Lab).
W2
What
Compresses, quantizes, and serves LLMs with high throughput.
W3
Where
Self-hosted on GPUs as a toolkit and server.
W4
When
When deploying open models efficiently in production.
W5
Why
Delivers fast serving with quantization and a persistent engine.
W6
With
GPUs (CUDA) and open model weights.
W7
Watch
Community and lab-maintained; open-source (Apache-2.0), self-hostable. Runs on your own GPU hardware; weights and traffic stay local.