FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness research paper by Stanford University, 2022
Stanford University · May 27, 2022 · Inference and efficiency · 5,331 citations · 16 upvotes
What it shows
FlashAttention computes exact attention with far fewer GPU memory reads and writes, making long sequences faster.
Summarised by hand from the abstract.
More from Stanford University
All 2Other inference and efficiency papers
TopicAbout this paper
- Authors
- Tri Dao, Daniel Y. Fu, Stefano Ermon and 2 more
- arXiv
- 2205.14135 · PDF
- Venue
- Neural Information Processing Systems
- Citations
- 5,331, 451 influential · Semantic Scholar
- Upvotes
- 16 · Hugging Face
- Lab
- Stanford University
Changes
| What changed | |
|---|---|
| Sep 24, 2026 | New paperAdded to the listSep 24, 2026 |
Sources: each lab's own papers and arXiv, with citation and upvote counts from Semantic Scholar and Hugging Face. One-line summaries are for orientation, not a substitute for the paper. Logos via logo.dev; trademarks belong to their owners.