Skip to content
Papers.

FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness research paper by Stanford University, 2022

Stanford University · May 27, 2022 · Inference and efficiency · 5,331 citations · 16 upvotes

Read on arXiv

What it shows

FlashAttention computes exact attention with far fewer GPU memory reads and writes, making long sequences faster.

Summarised by hand from the abstract.

More from Stanford University

All 2
Topic
About this paper
Authors
Tri Dao, Daniel Y. Fu, Stefano Ermon and 2 more
arXiv
2205.14135 · PDF
Venue
Neural Information Processing Systems
Citations
5,331, 451 influential · Semantic Scholar
Upvotes
16 · Hugging Face
Lab
Stanford University

Changes

What changed
New paperAdded to the listSep 24, 2026

Sources: each lab's own papers and arXiv, with citation and upvote counts from Semantic Scholar and Hugging Face. One-line summaries are for orientation, not a substitute for the paper. Logos via logo.dev; trademarks belong to their owners.

New papers by email

Monday afternoons, only in weeks with new papers from the labs.

Double opt-in. Unsubscribe any time.