← back to the directory
Application Inference & ServingAutomation & Integration

llama-swap

1Swaps 2local 3LLM 4models 5on 6demand

the six Ws · specification

W1 Who

Self hosters running multiple local models who need automatic swapping without manual server restarts.

W2 What

llama-swap is a lightweight reverse proxy that automatically starts, stops, and swaps backend LLM processes such as llama.cpp or vLLM based on the incoming request's model name.

W3 Where

Runs as a single Go binary or Docker container in front of one or more local inference servers.

W4 When

Actively maintained since 2024 with frequent releases.

W5 Why

Lets a single machine host many large models without keeping them all loaded in VRAM simultaneously.

W6 With

Wraps any OpenAI or Anthropic compatible backend, most commonly llama.cpp, vLLM, or TabbyAPI.

W7 Watch

Open source under MIT on GitHub, no account required, runs entirely on the operator's own hardware.

model-routingllama-cppproxyhot-swaplocal-llm

for agents & scripts

Reading this as a machine? Query it directly.

Search is open JSON - no key. Report telemetry after using a tool and it feeds that tool’s Proof Score. Or speak MCP to /mcp and discover tools mid-loop.