DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression research paper by DeepSeek, 2026
DeepSeek · Sep 17, 2026 · Inference and efficiency · 7 citations · 175 upvotes
What it shows
A 552B mixture-of-experts model with a 1M-token context, built to shrink the KV cache for agent workloads.
Summarised by hand from the abstract.
More from DeepSeek
All 2Releases that name it
From releases.fru.dev| Release | ||
|---|---|---|
| Sep 11, 2026 | DeepSeek V4.1 Flash is now available on Unity Gatewaydatabricks · Sep 11, 2026 | databricks |
| Sep 9, 2026 | DeepSeek V4.1 Flash now available on AI Gatewayvercel · Sep 9, 2026 | vercel |
Other inference and efficiency papers
TopicAbout this paper
- Authors
- DeepSeek-AI, Anyi Xu, B. Li and 589 more
- arXiv
- 2609.19969 · PDF
- Citations
- 7, 0 influential · Semantic Scholar
- Upvotes
- 175 · Hugging Face
- Lab
- DeepSeek
Changes
| What changed | |
|---|---|
| Sep 24, 2026 | New paperAdded to the listSep 24, 2026 |
Sources: each lab's own papers and arXiv, with citation and upvote counts from Semantic Scholar and Hugging Face. One-line summaries are for orientation, not a substitute for the paper. Logos via logo.dev; trademarks belong to their owners.