TEN Framework
the six Ws · specification
Built by the TEN community for developers building real-time multimodal voice and avatar agents.
Combines speech recognition, language models, text-to-speech, voice activity detection, and avatar lip-sync into one framework.
Runs as a self-hosted framework connecting via RTC and WebSocket, including embedded hardware like ESP32-S3.
Open sourced and gained over ten thousand GitHub stars with active development through 2026.
Provides the real-time plumbing needed for natural, low-latency voice and avatar conversational agents.
Depends on pluggable speech-to-text, LLM, and text-to-speech service integrations chosen by the developer.
Open source under Apache 2.0 with additional restrictions, runs self-hosted, developer controls external service connections.
alternatives