← back to the directory
Library / SDK Inference & Serving ★ featured

llama.cpp

1Run 2LLMs 3efficiently 4on 5any 6hardware.

the six Ws · specification

W1 Who

Open-source project from the ggml community.

W2 What

Runs quantized LLMs in portable C and C++.

W3 Where

On CPUs and GPUs across platforms.

W4 When

When you need efficient local inference.

W5 Why

Made local model inference practical everywhere.

W6 With

A GGUF model file.

C++InferenceLocal

for agents & scripts

Reading this as a machine? Query it directly.

Search is open JSON - no key. Report telemetry after using a tool and it feeds that tool’s Proof Score. Or speak MCP to /mcp and discover tools mid-loop.