← back to the directory
Library / SDK Inference & Serving

SGLang

1Fast 2serving 3for 4large 5language 6models

the six Ws · specification

W1 Who

Maintained by the SGLang project community (sgl-project).

W2 What

A fast serving framework for LLMs and VLMs with RadixAttention prefix caching and structured generation.

W3 Where

Runs self-hosted on GPUs as a Python serving runtime and API server.

W4 When

Reach for it when you need high-throughput inference with structured output or prefix reuse.

W5 Why

RadixAttention and efficient scheduling deliver strong throughput and latency.

W6 With

GPUs (CUDA) and model weights; integrates with Hugging Face models.

W7 Watch

Community-driven open-source (Apache-2.0); self-hostable and offline-capable; GPU hardware required for realistic performance.

InferenceServingLLMGPU

for agents & scripts

Reading this as a machine? Query it directly.

Search is open JSON - no key. Report telemetry after using a tool and it feeds that tool’s Proof Score. Or speak MCP to /mcp and discover tools mid-loop.