← back to the directory
Hosted Service Inference & Serving

DeepInfra

1Cheap 2serverless 3inference 4for 5open-source 6models

the six Ws · specification

W1 Who

Developers seeking the lowest-cost hosted inference for popular open-source LLMs.

W2 What

DeepInfra is a pay-as-you-go serverless API that hosts 50+ open-source LLM, image, and embedding models behind an OpenAI-compatible endpoint.

W3 Where

Accessed over HTTPS via DeepInfra's public API, with no infrastructure to manage.

W4 When

Operating since 2022 and continuously adding new open-weight model releases through 2026.

W5 Why

It exists to make running open-source models cheaper and simpler than self-hosting GPUs.

W6 With

Compatible with OpenAI SDK clients and frameworks like LangChain and LiteLLM.

W7 Watch

SaaS, proprietary, requires an API key; billed per token with free trial credits; requests and data are processed on DeepInfra's servers.

serverless inferencepay-per-tokenOpenAI-compatible APIopen-source modelsbatch inference

for agents & scripts

Reading this as a machine? Query it directly.

Search is open JSON - no key. Report telemetry after using a tool and it feeds that tool’s Proof Score. Or speak MCP to /mcp and discover tools mid-loop.