← back to the directory
Library / SDK Inference & ServingFine-tuning & Training

FlashAttention

1IO 2aware 3exact 4attention 5CUDA 6kernels

the six Ws · specification

W1 Who

Created by Tri Dao and collaborators at the Dao-AILab research group.

W2 What

FlashAttention is an IO aware exact attention algorithm that computes attention faster and with far less memory than naive implementations.

W3 Where

Hosted on GitHub at Dao-AILab/flash-attention and distributed via PyPI.

W4 When

First published in 2022, with FlashAttention-2 and FlashAttention-3 following in 2023 and 2024.

W5 Why

It removed the quadratic memory bottleneck of standard attention, enabling much longer context windows in transformer training and inference.

W6 With

Built as a CUDA extension for PyTorch and used inside Transformers, vLLM and many other frameworks.

W7 Watch

BSD 3-Clause licensed, open source, maintained by Tri Dao and Dao-AILab, roughly 25k GitHub stars.

attention-kernelgpu-optimizationmemory-efficientcudatransformers

for agents & scripts

Reading this as a machine? Query it directly.

Search is open JSON - no key. Report telemetry after using a tool and it feeds that tool’s Proof Score. Or speak MCP to /mcp and discover tools mid-loop.