← back to the directory
Hosted Service Inference & Serving

Groq

1Ultra-fast 2LLM 3inference 4on 5custom 6chips.

the six Ws · specification

W1 Who

By Groq.

W2 What

Serves open models at very low latency on custom LPU hardware.

W3 Where

A hosted API over HTTPS.

W4 When

When latency matters for inference.

W5 Why

Some of the fastest hosted token throughput.

W6 With

An API key and outbound HTTPS.

W7 Watch

Commercial API, API key; prompts are sent to Groq's cloud, billed per token.

LPULow-latencyAPI

for agents & scripts

Reading this as a machine? Query it directly.

Search is open JSON - no key. Report telemetry after using a tool and it feeds that tool’s Proof Score. Or speak MCP to /mcp and discover tools mid-loop.