Skip to content
Papers.

SwiftKV: Fast Prefill-Optimized Inference with Knowledge-Preserving Model Transformation research paper by Snowflake, 2024

Snowflake · Oct 4, 2024 · Inference and efficiency · 18 citations · 2 upvotes

Read on arXiv

What it shows

SwiftKV lets prompt tokens skip later layers and merges KV caches, cutting prefill cost for long-prompt workloads.

Summarised by hand from the abstract.

More from Snowflake

All 3
Topic
About this paper
Authors
Aurick Qiao, Zhewei Yao, Samyam Rajbhandari and 1 more
arXiv
2410.03960 · PDF
Venue
Conference on Empirical Methods in Natural Language Processing
Citations
18, 1 influential · Semantic Scholar
Upvotes
2 · Hugging Face
Lab
Snowflake · on Companies · on Acquisitions · on Quarterly · on Paydays · on Releases · on TechConf

Changes

What changed
New paperAdded to the listSep 24, 2026

Sources: each lab's own papers and arXiv, with citation and upvote counts from Semantic Scholar and Hugging Face. One-line summaries are for orientation, not a substitute for the paper. Logos via logo.dev; trademarks belong to their owners.

New papers by email

Monday afternoons, only in weeks with new papers from the labs.

Double opt-in. Unsubscribe any time.