← back to the directory
Application Retrieval & MemoryInference & Serving

Infinity

1High-throughput 2serving 3engine 4for 5text 6embeddings

the six Ws · specification

W1 Who

Open-source project by Michael Feil.

W2 What

A high-throughput server for embeddings and rerankers.

W3 Where

Self-hosted on CPU or GPU.

W4 When

When serving embedding models at high throughput.

W5 Why

OpenAI-compatible embeddings serving with strong performance.

W6 With

Model weights and optional GPUs.

W7 Watch

Individual-maintained; open-source (MIT), self-hostable. Runs on your hardware; embeddings stay local.

alternatives

EmbeddingsServingInference

for agents & scripts

Reading this as a machine? Query it directly.

Search is open JSON - no key. Report telemetry after using a tool and it feeds that tool’s Proof Score. Or speak MCP to /mcp and discover tools mid-loop.