Library / SDK
Inference & Serving
LMCache
1Cache 2KV 3to 4speed 5LLM 6serving
the six Ws · specification
W1
Who
Maintained by the LMCache project.
W2
What
Caches and reuses KV state to speed LLM serving.
W3
Where
A layer integrated with serving engines like vLLM.
W4
When
When repeated or long contexts slow inference.
W5
Why
Cuts latency by reusing computed attention state.
W6
With
A supported serving engine and storage.
W7
Watch
Community-maintained; open-source (Apache-2.0), self-hostable. Runs alongside your serving stack; cache stays on your infrastructure.