← back to the directory
Library / SDK Inference & Serving

GGML

1Low 2level 3C 4tensor 5inference 6library

the six Ws · specification

W1 Who

Created by Georgi Gerganov and maintained by the ggml-org community.

W2 What

GGML is a low level C tensor library for machine learning inference with built in quantization support.

W3 Where

Hosted on GitHub at ggml-org/ggml, migrated from the original ggerganov/ggml repository.

W4 When

Released in 2022 and now underpins llama.cpp and whisper.cpp.

W5 Why

It provides the lightweight, dependency free tensor and quantization foundation that made running large models on ordinary laptops practical.

W6 With

Has no external dependencies and offers optional CUDA, Metal and Vulkan backends.

W7 Watch

MIT licensed, open source, maintained by Georgi Gerganov and the ggml-org community, roughly 15k GitHub stars.

tensor-libraryc-languagequantizationcpu-inferenceon-device

for agents & scripts

Reading this as a machine? Query it directly.

Search is open JSON - no key. Report telemetry after using a tool and it feeds that tool’s Proof Score. Or speak MCP to /mcp and discover tools mid-loop.