SGLang
the six Ws · specification
Maintained by the SGLang project community (sgl-project).
A fast serving framework for LLMs and VLMs with RadixAttention prefix caching and structured generation.
Runs self-hosted on GPUs as a Python serving runtime and API server.
Reach for it when you need high-throughput inference with structured output or prefix reuse.
RadixAttention and efficient scheduling deliver strong throughput and latency.
GPUs (CUDA) and model weights; integrates with Hugging Face models.
Community-driven open-source (Apache-2.0); self-hostable and offline-capable; GPU hardware required for realistic performance.
works with