TabbyAPI
the six Ws · specification
Hobbyists and developers running quantized models locally who want a fast Exllama backed API server.
TabbyAPI is the official lightweight API server for the ExLlamaV2 inference backend, exposing an OpenAI compatible endpoint for GPTQ and EXL2 quantized models.
Self hosted on a local machine or cloud GPU instance, installed from source or via Docker.
Maintained continuously since 2023 by TheRoyalLab alongside the ExLlamaV2 project.
Offers very low latency single or multi user inference for consumer and prosumer GPUs running EXL2 quantized weights.
Requires the ExLlamaV2 Python library and a CUDA capable NVIDIA GPU.
Open source under AGPL 3.0 on GitHub, no signup needed, runs fully offline on the operator's own hardware.
alternatives
works with