← back to the directory
Library / SDK Fine-tuning & TrainingInference & Serving

LLM Compressor

1Quantizes 2LLMs 3for 4efficient 5vLLM 6deployment

the six Ws · specification

W1 Who

ML engineers preparing large language models for cheaper, faster deployment on vLLM.

W2 What

LLM Compressor is a library providing quantization algorithms like GPTQ, AWQ, and SmoothQuant for weight, activation, and KV cache compression.

W3 Where

Used as a Python library integrated into model preparation pipelines before deployment on vLLM inference servers.

W4 When

Developed by the vLLM project and Red Hat AI, actively maintained with releases through 2026.

W5 Why

Shrinks model memory footprint and speeds up inference by producing quantized checkpoints compatible with vLLM's serving engine.

W6 With

Depends on Hugging Face Transformers models as input and is designed to produce checkpoints consumed by vLLM.

W7 Watch

Open source under Apache-2.0, maintained by the vLLM project and Red Hat AI; runs entirely locally with no external API calls required.

quantizationmodel compressionGPTQvLLM integrationweight compression

for agents & scripts

Reading this as a machine? Query it directly.

Search is open JSON - no key. Report telemetry after using a tool and it feeds that tool’s Proof Score. Or speak MCP to /mcp and discover tools mid-loop.