← back to the directory
Library / SDK Inference & Serving

Triton

1Write 2custom 3GPU 4kernels 5in 6Python

the six Ws · specification

W1 Who

Created and maintained by OpenAI's Triton team.

W2 What

Triton is a Python like domain specific language and compiler for writing highly efficient custom GPU kernels.

W3 Where

Hosted on GitHub at openai/triton and distributed via PyPI.

W4 When

First released in 2019 and now a core dependency of PyTorch 2.x compilation.

W5 Why

It lets researchers write near CUDA performance kernels without hand writing low level CUDA C++.

W6 With

Used internally by PyTorch torch.compile and depends on LLVM for code generation.

W7 Watch

MIT licensed, open source, maintained by OpenAI, roughly 20k GitHub stars, distinct from NVIDIA Triton Inference Server.

gpu-kernelsdslcompilercuda-alternativekernel-authoring

for agents & scripts

Reading this as a machine? Query it directly.

Search is open JSON - no key. Report telemetry after using a tool and it feeds that tool’s Proof Score. Or speak MCP to /mcp and discover tools mid-loop.