← back to the directory
Model Inference & Serving

Moshi

1Real-time 2full-duplex 3spoken 4dialogue 5foundation 6model

the six Ws · specification

W1 Who

Kyutai, a French AI research lab, developed and released Moshi.

W2 What

Moshi is a full-duplex speech-to-speech foundation model that listens and speaks simultaneously for natural real-time conversation.

W3 Where

Hosted on GitHub at kyutai-labs/moshi and on Hugging Face under the kyutai organization.

W4 When

Released in September 2024.

W5 Why

It demonstrates low-latency, full-duplex voice interaction as an open alternative to proprietary real-time voice assistants.

W6 With

Runs on PyTorch or MLX with the Mimi neural audio codec bundled in the release.

W7 Watch

Released under CC-BY 4.0 for weights and MIT for code by Kyutai; fully open, no gating.

speech-to-speechfull-duplexreal-timeKyutai

for agents & scripts

Reading this as a machine? Query it directly.

Search is open JSON - no key. Report telemetry after using a tool and it feeds that tool’s Proof Score. Or speak MCP to /mcp and discover tools mid-loop.