Library / SDK
Inference & Serving
★ featured
llama.cpp
1Run 2LLMs 3efficiently 4on 5any 6hardware.
the six Ws · specification
W1
Who
Open-source project from the ggml community.
W2
What
Runs quantized LLMs in portable C and C++.
W3
Where
On CPUs and GPUs across platforms.
W4
When
When you need efficient local inference.
W5
Why
Made local model inference practical everywhere.
W6
With
A GGUF model file.
works with