Efficient Memory Management for Large Language Model Serving with PagedAttention research paper by UC Berkeley, 2023
UC Berkeley · Sep 12, 2023 · Inference and efficiency · 8,469 citations · 72 upvotes
What it shows
vLLM manages the KV cache like pages of virtual memory, serving models with 2 to 4 times the throughput.
Summarised by hand from the abstract.
More from UC Berkeley
All 2Other inference and efficiency papers
TopicAbout this paper
- Authors
- Woosuk Kwon, Zhuohan Li, Siyuan Zhuang and 6 more
- arXiv
- 2309.06180 · PDF
- Venue
- Symposium on Operating Systems Principles
- Citations
- 8,469, 1,165 influential · Semantic Scholar
- Upvotes
- 72 · Hugging Face
- Code
- github.com/vllm-project/vllm
- Lab
- UC Berkeley
Changes
| What changed | |
|---|---|
| Sep 24, 2026 | New paperAdded to the listSep 24, 2026 |
Sources: each lab's own papers and arXiv, with citation and upvote counts from Semantic Scholar and Hugging Face. One-line summaries are for orientation, not a substitute for the paper. Logos via logo.dev; trademarks belong to their owners.