Library / SDK
Inference & Serving
MLC LLM
1Run 2LLMs 3natively 4on 5any 6device
the six Ws · specification
W1
Who
Maintained by the MLC AI community.
W2
What
Compiles and runs LLMs natively across devices and GPUs.
W3
Where
On phones, browsers, and desktops via a compiled runtime.
W4
When
When you need local inference on diverse hardware.
W5
Why
Universal deployment through machine-learning compilation.
W6
With
The MLC runtime and compiled model weights.
W7
Watch
Community-maintained; open-source (Apache-2.0). Runs fully on-device, so inference stays local and offline-capable.