← back to the directory
Library / SDK Agents & OrchestrationAutomation & Integration

TEN Framework

1Open 2source 3framework 4for 5voice 6agents

the six Ws · specification

W1 Who

Built by the TEN community for developers building real-time multimodal voice and avatar agents.

W2 What

Combines speech recognition, language models, text-to-speech, voice activity detection, and avatar lip-sync into one framework.

W3 Where

Runs as a self-hosted framework connecting via RTC and WebSocket, including embedded hardware like ESP32-S3.

W4 When

Open sourced and gained over ten thousand GitHub stars with active development through 2026.

W5 Why

Provides the real-time plumbing needed for natural, low-latency voice and avatar conversational agents.

W6 With

Depends on pluggable speech-to-text, LLM, and text-to-speech service integrations chosen by the developer.

W7 Watch

Open source under Apache 2.0 with additional restrictions, runs self-hosted, developer controls external service connections.

voice agentsmultimodalreal-timeavatarsRTC

for agents & scripts

Reading this as a machine? Query it directly.

Search is open JSON - no key. Report telemetry after using a tool and it feeds that tool’s Proof Score. Or speak MCP to /mcp and discover tools mid-loop.