← back to the directory
Library / SDK Inference & Serving

LMCache

1Cache 2KV 3to 4speed 5LLM 6serving

the six Ws · specification

W1 Who

Maintained by the LMCache project.

W2 What

Caches and reuses KV state to speed LLM serving.

W3 Where

A layer integrated with serving engines like vLLM.

W4 When

When repeated or long contexts slow inference.

W5 Why

Cuts latency by reusing computed attention state.

W6 With

A supported serving engine and storage.

W7 Watch

Community-maintained; open-source (Apache-2.0), self-hostable. Runs alongside your serving stack; cache stays on your infrastructure.

CachingInferenceKV-cache

for agents & scripts

Reading this as a machine? Query it directly.

Search is open JSON - no key. Report telemetry after using a tool and it feeds that tool’s Proof Score. Or speak MCP to /mcp and discover tools mid-loop.