Modular MAX
the six Ws · specification
ML infra teams wanting a high-performance, hardware-portable alternative to standard inference stacks.
Modular MAX is a graph-compiled inference engine and serving platform, built on the Mojo language, that runs LLMs across NVIDIA, AMD, and Apple silicon GPUs.
Self-hosted, deployed via Docker containers or the Modular CLI on GPU servers or cloud instances.
Under active development since 2023; Qualcomm announced its acquisition of Modular in June 2026.
It exists to give teams a single, portable, high-throughput inference runtime independent of vendor-specific stacks.
Ships with the Mojo standard library and a large kernel library targeting multiple GPU vendors.
Open source core (Modular Platform) with a commercial layer; self-hosted; no data sent externally unless using Modular's hosted offerings.
alternatives