Skip to content
Papers.

Efficient Memory Management for Large Language Model Serving with PagedAttention research paper by UC Berkeley, 2023

UC Berkeley · Sep 12, 2023 · Inference and efficiency · 8,469 citations · 72 upvotes

Read on arXiv

What it shows

vLLM manages the KV cache like pages of virtual memory, serving models with 2 to 4 times the throughput.

Summarised by hand from the abstract.

More from UC Berkeley

All 2
Topic
About this paper
Authors
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang and 6 more
arXiv
2309.06180 · PDF
Venue
Symposium on Operating Systems Principles
Citations
8,469, 1,165 influential · Semantic Scholar
Upvotes
72 · Hugging Face
Code
github.com/vllm-project/vllm
Lab
UC Berkeley

Changes

What changed
New paperAdded to the listSep 24, 2026

Sources: each lab's own papers and arXiv, with citation and upvote counts from Semantic Scholar and Hugging Face. One-line summaries are for orientation, not a substitute for the paper. Logos via logo.dev; trademarks belong to their owners.

New papers by email

Monday afternoons, only in weeks with new papers from the labs.

Double opt-in. Unsubscribe any time.