Skip to content
Papers.

DualPath: Breaking the Storage Bandwidth Bottleneck in Agentic LLM Inference research paper by DeepSeek, 2026

DeepSeek · Feb 25, 2026 · Agents and evaluation · 19 citations · 55 upvotes · unverified

Read on arXiv

What it shows

DualPath addresses KV-cache storage I/O bottlenecks in multi-turn LLM inference by introducing dual-path loading and dynamic load balancing across prefill and decode engines.

UnverifiedHugging Face's summary; not yet checked by hand.

More from DeepSeek

All 22
PaperCitations
DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache CompressionA 552B mixture-of-experts model with a 1M-token context, built to shrink the KV cache for agent workloads.Inference and efficiency · Sep 20267
DSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive GenerationDSpark enhances LLM inference speed by combining parallel draft generation with adaptive verification that reduces waste and improves throughput in high-concurrency settings.Inference and efficiency · Jul 2026 · Unverified30
DeepSeek-OCR 2: Visual Causal FlowDeepSeek-OCR 2 introduces DeepEncoder V2 that dynamically reorders visual tokens based on semantic content, enabling more human-like causal reasoning in 2D image understanding through cascaded 1D causal structures.Multimodal and robotics · Jan 2026 · Unverified75
mHC: Manifold-Constrained Hyper-ConnectionsManifold-Constrained Hyper-Connections (mHC) stabilize and scale residual connection architectures by restoring identity mapping properties through manifold projection and infrastructure optimization.Training and scaling · Dec 2025 · Unverified82
DeepSeek-V3.2: Pushing the Frontier of Open Large Language ModelsDeepSeek-V3.2 introduces DeepSeek Sparse Attention and a scalable reinforcement learning framework, achieving superior reasoning and performance compared to GPT-5 and Gemini-3.0-Pro in complex reasoning tasks.Agents and evaluation · Dec 2025 · Unverified761
DeepSeekMath-V2: Towards Self-Verifiable Mathematical ReasoningA self-verifying large language model for theorem proving improves mathematical reasoning by incentivizing rigorous step-by-step derivations and achieves high scores in international competitions.Reasoning · Nov 2025 · Unverified68
DeepSeek-OCR: Contexts Optical CompressionDeepSeek-OCR uses optical 2D mapping to compress long contexts, achieving high OCR precision with reduced vision tokens and demonstrating practical value in document processing.Inference and efficiency · Oct 2025 · Unverified189
Insights into DeepSeek-V3: Scaling Challenges and Reflections on Hardware for AI ArchitecturesDeepSeek-V3 addresses hardware limitations through MLA, MoE, FP8 training, and Multi-Plane Network Topology, enabling efficient large-scale LLM training and inference.Inference and efficiency · May 2025 · Unverified107
Topic
PaperCitations
Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic CapabilitiesGemini 2.X model family, including Gemini 2.5 Pro and Flash, offers superior coding, reasoning, and multimodal understanding capabilities across a range of computational efficiencies.Google · Jul 2025 · Unverified4,134
WebWatcher: Breaking New Frontier of Vision-Language Deep Research AgentWebWatcher, a multimodal agent with enhanced visual-language reasoning, outperforms existing agents in complex visual and textual information retrieval tasks using synthetic trajectories and reinforcement learning.Alibaba (Qwen) · Aug 2025 · Unverified117
SkillOpt: Executive Strategy for Self-Evolving Agent SkillsSkillOpt introduces a systematic text-space optimizer for agent skills that trains skills as external agent state with stable updates and zero deployment inference overhead, achieving superior performance across multiple benchmarks and execution environments.Microsoft · May 2026 · Unverified85
Nemotron 3 Nano: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic ReasoningWe present Nemotron 3 Nano 30B-A3B, a Mixture-of-Experts hybrid Mamba-Transformer language model.NVIDIA · Dec 2025 · Unverified81
AgentFold: Long-Horizon Web Agents with Proactive Context ManagementAgentFold, a novel proactive context management paradigm, enhances long-horizon task performance through dynamic context folding, achieving superior results on benchmarks compared to larger models and proprietary agents.Alibaba (Qwen) · Oct 2025 · Unverified77
Agent Learning via Early ExperienceEarly experience, using agent-generated interaction data without reward signals, improves policy effectiveness and generalization, serving as a bridge between imitation learning and reinforcement learning.Meta · Oct 2025 · Unverified64
About this paper
Authors
Yongtong Wu, Shaoyuan Chen, Yinmin Zhong and 10 more
arXiv
2602.21548 · PDF
Venue
arXiv.org
Citations
19, 2 influential · Semantic Scholar
Upvotes
55 · Hugging Face
Lab
DeepSeek · on Companies

Changes

What changed
Influential citationsfirst count: 2Sep 25, 2026
Citationsfirst count: 19Sep 25, 2026
New paperFound by the weekly scan, unverifiedSep 25, 2026

Sources: each lab's own papers and arXiv, with citation and upvote counts from Semantic Scholar and Hugging Face. One-line summaries are for orientation, not a substitute for the paper. Logos via logo.dev; trademarks belong to their owners.

New papers by email

Monday afternoons, only in weeks with new papers from the labs.

Double opt-in. Unsubscribe any time.