SwiftKV: Fast Prefill-Optimized Inference with Knowledge-Preserving Model Transformation research paper by Snowflake, 2024
Snowflake · Oct 4, 2024 · Inference and efficiency · 18 citations · 2 upvotes
What it shows
SwiftKV lets prompt tokens skip later layers and merges KV caches, cutting prefill cost for long-prompt workloads.
Summarised by hand from the abstract.
More from Snowflake
All 3Other inference and efficiency papers
TopicAbout this paper
- Authors
- Aurick Qiao, Zhewei Yao, Samyam Rajbhandari and 1 more
- arXiv
- 2410.03960 · PDF
- Venue
- Conference on Empirical Methods in Natural Language Processing
- Citations
- 18, 1 influential · Semantic Scholar
- Upvotes
- 2 · Hugging Face
- Lab
- Snowflake · on Companies · on Acquisitions · on Quarterly · on Paydays · on Releases · on TechConf
Changes
| What changed | |
|---|---|
| Sep 24, 2026 | New paperAdded to the listSep 24, 2026 |
Sources: each lab's own papers and arXiv, with citation and upvote counts from Semantic Scholar and Hugging Face. One-line summaries are for orientation, not a substitute for the paper. Logos via logo.dev; trademarks belong to their owners.