Library / SDK
Inference & Serving
ExLlamaV2
1Fast 2quantized 3LLM 4inference 5on 6GPUs.
the six Ws · specification
W1
Who
By turboderp.
W2
What
A fast inference library for running quantized models on consumer GPUs.
W3
Where
A Python/C++ library on GPUs.
W4
When
When squeezing models onto local GPUs.
W5
Why
High-speed inference for quantized weights.
W6
With
Python, CUDA, and a GPU.
W7
Watch
MIT open-source; runs locally, no credentials.
alternatives
works with
QuantizationGPUFast