← back to the directory
Model Inference & Serving

F5-TTS

1Zero-shot 2voice 3cloning 4via 5flow 6matching

the six Ws · specification

W1 Who

Researchers and developers needing high fidelity zero-shot voice cloning from short audio samples.

W2 What

An open weights text-to-speech model that clones a voice from a few seconds of reference audio using flow matching.

W3 Where

Self-hosted inference via Python and a Gradio demo, or served through the Hugging Face model hub.

W4 When

Published in 2024 with active updates and community adoption continuing through 2026.

W5 Why

Delivers fluent, faithful zero-shot voice cloning quality that rivals closed commercial TTS systems.

W6 With

Requires Python, PyTorch and a CUDA GPU for practical inference speed.

W7 Watch

Open source under MIT with open model weights, self-hosted, no credentials, over 15k GitHub stars.

voice-cloningzero-shot-ttsflow-matchingtext-to-speechopen-weights

for agents & scripts

Reading this as a machine? Query it directly.

Search is open JSON - no key. Report telemetry after using a tool and it feeds that tool’s Proof Score. Or speak MCP to /mcp and discover tools mid-loop.