Library / SDK
Inference & Serving
★ featured
vLLM
1Fast, 2high-throughput 3inference 4and 5serving 6engine.
the six Ws · specification
W1
Who
Open-source inference engine from the vLLM project.
W2
What
Serves open LLMs with high throughput and efficiency.
W3
Where
Runs on your own GPU servers.
W4
When
When self-hosting models at scale.
W5
Why
PagedAttention delivers fast, memory-efficient serving.
W6
With
GPUs and open model weights.
alternatives
works with
- Aya Expanse
- Beam
- CodeQwen1.5
- Command R
- DBRX
- DeepSeek
- DeepSeek-Coder-V2
- FlashAttention
- GLM-4
- GPTQModel
- Granite
- Jamba 1.5
- Kimi K2
- LiteLLM
- Llama
- Llama Stack
- Microsoft AICI
- Mistral
- Mixtral 8x7B
- Modal
- Molmo
- OLMo 2
- OpenRLHF
- Pixtral 12B
- Qwen
- Qwen2-VL
- Ray
- RunPod
- Snowflake Arctic
- StarCoder2
- Transformers
- Triton
- XGrammar
- Yi
- gpt-oss
- lm-evaluation-harness
- lm-format-enforcer
- verl