Skip to content
Papers.

DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression research paper by DeepSeek, 2026

DeepSeek · Sep 17, 2026 · Inference and efficiency · 7 citations · 175 upvotes

Read on arXiv

What it shows

A 552B mixture-of-experts model with a 1M-token context, built to shrink the KV cache for agent workloads.

Summarised by hand from the abstract.

More from DeepSeek

All 2

Releases that name it

From releases.fru.dev
Release
DeepSeek V4.1 Flash is now available on Unity Gatewaydatabricks · Sep 11, 2026
DeepSeek V4.1 Flash now available on AI Gatewayvercel · Sep 9, 2026
Topic
About this paper
Authors
DeepSeek-AI, Anyi Xu, B. Li and 589 more
arXiv
2609.19969 · PDF
Citations
7, 0 influential · Semantic Scholar
Upvotes
175 · Hugging Face
Lab
DeepSeek

Changes

What changed
New paperAdded to the listSep 24, 2026

Sources: each lab's own papers and arXiv, with citation and upvote counts from Semantic Scholar and Hugging Face. One-line summaries are for orientation, not a substitute for the paper. Logos via logo.dev; trademarks belong to their owners.

New papers by email

Monday afternoons, only in weeks with new papers from the labs.

Double opt-in. Unsubscribe any time.