← back to the directory
Library / SDK Inference & Serving

Text Generation Inference

1Production 2LLM 3serving 4from 5Hugging 6Face

the six Ws · specification

W1 Who

Maintained by Hugging Face.

W2 What

A toolkit for deploying and serving LLMs with high-throughput, low-latency text generation over HTTP and gRPC.

W3 Where

Runs self-hosted, primarily on GPUs, as a Rust/Python server and Docker image.

W4 When

Reach for it when serving open LLMs in production with continuous batching and streaming.

W5 Why

Optimized inference purpose-built for the Hugging Face model ecosystem.

W6 With

GPUs (CUDA) for best performance and model weights from the Hugging Face Hub.

W7 Watch

Built by Hugging Face; open-source (Apache-2.0 in current releases); self-hostable and runs offline once weights are cached; GPU strongly recommended.

InferenceServingLLMGPU

for agents & scripts

Reading this as a machine? Query it directly.

Search is open JSON - no key. Report telemetry after using a tool and it feeds that tool’s Proof Score. Or speak MCP to /mcp and discover tools mid-loop.