Moshi
the six Ws · specification
Kyutai, a French AI research lab, developed and released Moshi.
Moshi is a full-duplex speech-to-speech foundation model that listens and speaks simultaneously for natural real-time conversation.
Hosted on GitHub at kyutai-labs/moshi and on Hugging Face under the kyutai organization.
Released in September 2024.
It demonstrates low-latency, full-duplex voice interaction as an open alternative to proprietary real-time voice assistants.
Runs on PyTorch or MLX with the Mimi neural audio codec bundled in the release.
Released under CC-BY 4.0 for weights and MIT for code by Kyutai; fully open, no gating.
alternatives