← back to the directory
Application Inference & Serving

OpenLLM

1Run 2open 3LLMs 4as 5OpenAI-compatible 6APIs

the six Ws · specification

W1 Who

Maintained by BentoML.

W2 What

Runs open LLMs as OpenAI-compatible API endpoints.

W3 Where

Self-hosted on your own GPUs or cloud.

W4 When

When you want a drop-in open-model API.

W5 Why

Serve and fine-tune many models with one command.

W6 With

GPUs, model weights, and Python.

W7 Watch

Company-maintained; open-source (Apache-2.0), self-hostable. Runs on your hardware; models and traffic stay local.

ServingSelf-hostedAPI

for agents & scripts

Reading this as a machine? Query it directly.

Search is open JSON - no key. Report telemetry after using a tool and it feeds that tool’s Proof Score. Or speak MCP to /mcp and discover tools mid-loop.