Hosted Service
Inference & Serving
Groq
1Ultra-fast 2LLM 3inference 4on 5custom 6chips.
the six Ws · specification
W1
Who
By Groq.
W2
What
Serves open models at very low latency on custom LPU hardware.
W3
Where
A hosted API over HTTPS.
W4
When
When latency matters for inference.
W5
Why
Some of the fastest hosted token throughput.
W6
With
An API key and outbound HTTPS.
W7
Watch
Commercial API, API key; prompts are sent to Groq's cloud, billed per token.
alternatives
works with
LPULow-latencyAPI