← back to the directory
Library / SDK Inference & Serving

Modular MAX

1Graph-compiled 2inference 3engine 4spans 5multiple 6hardware

the six Ws · specification

W1 Who

ML infra teams wanting a high-performance, hardware-portable alternative to standard inference stacks.

W2 What

Modular MAX is a graph-compiled inference engine and serving platform, built on the Mojo language, that runs LLMs across NVIDIA, AMD, and Apple silicon GPUs.

W3 Where

Self-hosted, deployed via Docker containers or the Modular CLI on GPU servers or cloud instances.

W4 When

Under active development since 2023; Qualcomm announced its acquisition of Modular in June 2026.

W5 Why

It exists to give teams a single, portable, high-throughput inference runtime independent of vendor-specific stacks.

W6 With

Ships with the Mojo standard library and a large kernel library targeting multiple GPU vendors.

W7 Watch

Open source core (Modular Platform) with a commercial layer; self-hosted; no data sent externally unless using Modular's hosted offerings.

graph-compiled engineMojo kernelsmulti-hardwareGPU CUDA ROCm

for agents & scripts

Reading this as a machine? Query it directly.

Search is open JSON - no key. Report telemetry after using a tool and it feeds that tool’s Proof Score. Or speak MCP to /mcp and discover tools mid-loop.