← back to the directory
Library / SDK Inference & Serving

LMDeploy

1Compress, 2deploy, 3and 4serve 5LLMs 6efficiently

the six Ws · specification

W1 Who

Maintained by the InternLM team (Shanghai AI Lab).

W2 What

Compresses, quantizes, and serves LLMs with high throughput.

W3 Where

Self-hosted on GPUs as a toolkit and server.

W4 When

When deploying open models efficiently in production.

W5 Why

Delivers fast serving with quantization and a persistent engine.

W6 With

GPUs (CUDA) and open model weights.

W7 Watch

Community and lab-maintained; open-source (Apache-2.0), self-hostable. Runs on your own GPU hardware; weights and traffic stay local.

InferenceServingGPU

for agents & scripts

Reading this as a machine? Query it directly.

Search is open JSON - no key. Report telemetry after using a tool and it feeds that tool’s Proof Score. Or speak MCP to /mcp and discover tools mid-loop.