← back to the directory
Hosted Service Inference & Serving

FriendliAI

1Fast 2pay-per-token 3cloud 4inference 5for 6LLMs

the six Ws · specification

W1 Who

AI teams and enterprises needing production LLM inference without managing GPU infrastructure.

W2 What

FriendliAI is a hosted inference cloud offering serverless model APIs, dedicated GPU endpoints, and containerized deployment for open and custom LLMs.

W3 Where

Accessed via FriendliAI's cloud API or deployed into a customer's own VPC (BYOC) for enterprise plans.

W4 When

Commercially operating since 2023, with ongoing 2026 feature additions such as InferenceSense for idle GPU monetization.

W5 Why

It exists to cut inference latency and cost for teams serving open-source and fine-tuned LLMs at scale.

W6 With

Works with Hugging Face model checkpoints and integrates via an OpenAI-compatible API.

W7 Watch

SaaS, proprietary, requires an API key; usage-based billing with enterprise BYOC option; prompts and data go to FriendliAI's cloud unless deployed on customer VPC.

LLM inferenceserverless endpointsdedicated GPUspay-per-tokenenterprise BYOC

for agents & scripts

Reading this as a machine? Query it directly.

Search is open JSON - no key. Report telemetry after using a tool and it feeds that tool’s Proof Score. Or speak MCP to /mcp and discover tools mid-loop.