← back to the directory
Model Inference & Serving

Sesame CSM

1Conversational 2speech 3model 4for 5natural 6dialogue

the six Ws · specification

W1 Who

Sesame AI Labs released the Conversational Speech Model, CSM.

W2 What

CSM is a speech generation model built on a Llama backbone that produces contextually natural, conversational sounding speech from text and prior audio context.

W3 Where

Hosted on GitHub at SesameAILabs/csm and on Hugging Face under sesame.

W4 When

Released in February 2025.

W5 Why

It targets the voice presence problem, aiming to make synthetic conversational speech sound less flat than typical TTS.

W6 With

Built on a Llama style backbone plus the Mimi audio codec, runnable via PyTorch.

W7 Watch

Released under the Apache 2.0 license by Sesame; base model open, though the full production voice demo is not fully released.

text-to-speechconversationalvoice-cloningSesame AI

for agents & scripts

Reading this as a machine? Query it directly.

Search is open JSON - no key. Report telemetry after using a tool and it feeds that tool’s Proof Score. Or speak MCP to /mcp and discover tools mid-loop.