← back to the directory
Library / SDK Inference & ServingRetrieval & Memory

Text Embeddings Inference

1Fast 2serving 3for 4text 5embedding 6models

the six Ws · specification

W1 Who

Maintained by Hugging Face.

W2 What

A fast server for embedding and reranking models.

W3 Where

Self-hosted on CPU or GPU.

W4 When

When serving embeddings at high throughput.

W5 Why

Optimized serving for the embedding half of RAG.

W6 With

Model weights and optional GPUs.

W7 Watch

Company-maintained (Hugging Face); open-source (Apache-2.0), self-hostable. Runs on your hardware; embeddings stay local.

EmbeddingsServingGPU

for agents & scripts

Reading this as a machine? Query it directly.

Search is open JSON - no key. Report telemetry after using a tool and it feeds that tool’s Proof Score. Or speak MCP to /mcp and discover tools mid-loop.