← back to the directory
Library / SDK Inference & Serving

MLC LLM

1Run 2LLMs 3natively 4on 5any 6device

the six Ws · specification

W1 Who

Maintained by the MLC AI community.

W2 What

Compiles and runs LLMs natively across devices and GPUs.

W3 Where

On phones, browsers, and desktops via a compiled runtime.

W4 When

When you need local inference on diverse hardware.

W5 Why

Universal deployment through machine-learning compilation.

W6 With

The MLC runtime and compiled model weights.

W7 Watch

Community-maintained; open-source (Apache-2.0). Runs fully on-device, so inference stays local and offline-capable.

InferenceOn-deviceCompilation

for agents & scripts

Reading this as a machine? Query it directly.

Search is open JSON - no key. Report telemetry after using a tool and it feeds that tool’s Proof Score. Or speak MCP to /mcp and discover tools mid-loop.