← back to the directory
Library / SDK Inference & Serving

Aphrodite Engine

1High 2throughput 3LLM 4serving 5engine 6fork

the six Ws · specification

W1 Who

Developers and researchers who need high throughput multi user LLM inference on their own GPUs.

W2 What

A vLLM derived inference engine, renamed Sonar by dphnAI, formerly PygmalionAI, that serves large language models with continuous batching, paged attention, and broad quantization support.

W3 Where

Self hosted on local or cloud GPUs via pip install or Docker, exposing an OpenAI compatible API.

W4 When

Originally released in 2023 as Aphrodite Engine and actively developed through 2026 under its new Sonar name.

W5 Why

Gives teams vLLM class throughput plus extra quantization and sampling options favored by the open model hosting community.

W6 With

Built on PyTorch and CUDA, and shares much of its core architecture with the vLLM project.

W7 Watch

Open source under AGPL 3.0 on GitHub, no account required, runs entirely on hardware the operator controls, the project renamed from Aphrodite Engine to Sonar and moved from the PygmalionAI to the dphnAI GitHub org.

llm-inferencevllm-forkquantizationpaged-attentiongpu-serving

for agents & scripts

Reading this as a machine? Query it directly.

Search is open JSON - no key. Report telemetry after using a tool and it feeds that tool’s Proof Score. Or speak MCP to /mcp and discover tools mid-loop.