Skip to content
Papers.

DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence research paper by DeepSeek, 2024

DeepSeek · Jun 17, 2024 · Reasoning · 502 citations · 71 upvotes · unverified

Read on arXiv

What it shows

DeepSeek-Coder-V2, a Mixture-of-Experts language model, excels in code-specific tasks by enhancing coding and mathematical reasoning capabilities while expanding language support and context length.

UnverifiedHugging Face's summary; not yet checked by hand.

More from DeepSeek

All 22
PaperCitations
DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache CompressionA 552B mixture-of-experts model with a 1M-token context, built to shrink the KV cache for agent workloads.Inference and efficiency · Sep 20267
DSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive GenerationDSpark enhances LLM inference speed by combining parallel draft generation with adaptive verification that reduces waste and improves throughput in high-concurrency settings.Inference and efficiency · Jul 2026 · Unverified30
DualPath: Breaking the Storage Bandwidth Bottleneck in Agentic LLM InferenceDualPath addresses KV-cache storage I/O bottlenecks in multi-turn LLM inference by introducing dual-path loading and dynamic load balancing across prefill and decode engines.Agents and evaluation · Feb 2026 · Unverified19
DeepSeek-OCR 2: Visual Causal FlowDeepSeek-OCR 2 introduces DeepEncoder V2 that dynamically reorders visual tokens based on semantic content, enabling more human-like causal reasoning in 2D image understanding through cascaded 1D causal structures.Multimodal and robotics · Jan 2026 · Unverified75
mHC: Manifold-Constrained Hyper-ConnectionsManifold-Constrained Hyper-Connections (mHC) stabilize and scale residual connection architectures by restoring identity mapping properties through manifold projection and infrastructure optimization.Training and scaling · Dec 2025 · Unverified82
DeepSeek-V3.2: Pushing the Frontier of Open Large Language ModelsDeepSeek-V3.2 introduces DeepSeek Sparse Attention and a scalable reinforcement learning framework, achieving superior reasoning and performance compared to GPT-5 and Gemini-3.0-Pro in complex reasoning tasks.Agents and evaluation · Dec 2025 · Unverified761
DeepSeekMath-V2: Towards Self-Verifiable Mathematical ReasoningA self-verifying large language model for theorem proving improves mathematical reasoning by incentivizing rigorous step-by-step derivations and achieves high scores in international competitions.Reasoning · Nov 2025 · Unverified68
DeepSeek-OCR: Contexts Optical CompressionDeepSeek-OCR uses optical 2D mapping to compress long contexts, achieving high OCR precision with reduced vision tokens and demonstrating practical value in document processing.Inference and efficiency · Oct 2025 · Unverified189
Topic
PaperCitations
Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsAsking a model to write out its intermediate steps (chain of thought) sharply improves its math and logic answers.Google · Jan 202222.1k
MolmoAct: Action Reasoning Models that can Reason in SpaceAction Reasoning Models (ARMs) integrate perception, planning, and control to enable adaptable and explainable robotic behavior, achieving superior performance across various tasks and settings.Ai2 · Aug 2025 · Unverified185
ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent PlanningThinkAct, a dual-system framework, uses reinforced visual latent planning to enable few-shot adaptation, long-horizon planning, and self-correction in embodied AI tasks by bridging high-level reasoning with low-level action execution.NVIDIA · Jul 2025 · Unverified172
WebShaper: Agentically Data Synthesizing via Information-Seeking FormalizationWebShaper, a formalization-driven framework, synthesizes information-seeking datasets using set theory and Knowledge Projections to enhance reasoning structure and achieve top performance in open-sourced benchmarks.Alibaba (Qwen) · Jul 2025 · Unverified113
Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?Self-distillation in large language models can degrade mathematical reasoning performance by suppressing uncertainty expression, particularly affecting out-of-distribution tasks.Microsoft · Mar 2026 · Unverified79
Soft Adaptive Policy OptimizationSoft Adaptive Policy Optimization (SAPO) enhances the stability and performance of reinforcement learning in large language models by adaptively attenuating off-policy updates with a smooth, temperature-controlled gate, leading to improved training stability and performance.Alibaba (Qwen) · Nov 2025 · Unverified79
About this paper
Authors
DeepSeek-AI, Qihao Zhu, Daya Guo and 37 more
arXiv
2406.11931 · PDF
Venue
arXiv.org
Citations
502, 47 influential · Semantic Scholar
Upvotes
71 · Hugging Face
Code
github.com/deepseek-ai/deepseek-coder-v2
Lab
DeepSeek · on Companies

Changes

What changed
Influential citationsfirst count: 47Sep 25, 2026
Citationsfirst count: 502Sep 25, 2026
New paperFound by the weekly scan, unverifiedSep 25, 2026

Sources: each lab's own papers and arXiv, with citation and upvote counts from Semantic Scholar and Hugging Face. One-line summaries are for orientation, not a substitute for the paper. Logos via logo.dev; trademarks belong to their owners.

New papers by email

Monday afternoons, only in weeks with new papers from the labs.

Double opt-in. Unsubscribe any time.