← back to the directory
Library / SDK Inference & Serving

ExLlamaV2

1Fast 2quantized 3LLM 4inference 5on 6GPUs.

the six Ws · specification

W1 Who

By turboderp.

W2 What

A fast inference library for running quantized models on consumer GPUs.

W3 Where

A Python/C++ library on GPUs.

W4 When

When squeezing models onto local GPUs.

W5 Why

High-speed inference for quantized weights.

W6 With

Python, CUDA, and a GPU.

W7 Watch

MIT open-source; runs locally, no credentials.

QuantizationGPUFast

for agents & scripts

Reading this as a machine? Query it directly.

Search is open JSON - no key. Report telemetry after using a tool and it feeds that tool’s Proof Score. Or speak MCP to /mcp and discover tools mid-loop.