← back to the directory
Library / SDK Inference & Serving

TabbyAPI

1Lightweight 2OpenAI 3compatible 4server 5for 6ExLlamaV2

the six Ws · specification

W1 Who

Hobbyists and developers running quantized models locally who want a fast Exllama backed API server.

W2 What

TabbyAPI is the official lightweight API server for the ExLlamaV2 inference backend, exposing an OpenAI compatible endpoint for GPTQ and EXL2 quantized models.

W3 Where

Self hosted on a local machine or cloud GPU instance, installed from source or via Docker.

W4 When

Maintained continuously since 2023 by TheRoyalLab alongside the ExLlamaV2 project.

W5 Why

Offers very low latency single or multi user inference for consumer and prosumer GPUs running EXL2 quantized weights.

W6 With

Requires the ExLlamaV2 Python library and a CUDA capable NVIDIA GPU.

W7 Watch

Open source under AGPL 3.0 on GitHub, no signup needed, runs fully offline on the operator's own hardware.

exllamav2gpu-inferenceopenai-compatiblelocal-llm

for agents & scripts

Reading this as a machine? Query it directly.

Search is open JSON - no key. Report telemetry after using a tool and it feeds that tool’s Proof Score. Or speak MCP to /mcp and discover tools mid-loop.