← back to the directory
Library / SDK Inference & Serving

GPTQModel

1Production 2grade 3GPTQ 4quantization 5for 6LLMs

the six Ws · specification

W1 Who

Maintained by ModelCloud as the successor to the original AutoGPTQ project.

W2 What

GPTQModel implements GPTQ post training quantization to compress large language model weights to 4-bit or lower with minimal accuracy loss.

W3 Where

Hosted on GitHub at ModelCloud/GPTQModel and distributed via PyPI.

W4 When

Released in 2024, having fully supplanted the now deprecated AutoGPTQ project.

W5 Why

It gives teams a maintained, faster path to quantized LLM weights after the original AutoGPTQ project stalled.

W6 With

Integrates with Hugging Face Transformers, Optimum and PEFT.

W7 Watch

Apache 2.0 licensed with some AGPL 3.0 kernels, open source, maintained by ModelCloud, roughly 1.2k GitHub stars, note that AutoGPTQ is deprecated in its favor.

gptqquantizationweight-compressionllm-inferenceautogptq-successor

for agents & scripts

Reading this as a machine? Query it directly.

Search is open JSON - no key. Report telemetry after using a tool and it feeds that tool’s Proof Score. Or speak MCP to /mcp and discover tools mid-loop.