Skip to content
Papers.

DeepSeek-OCR: Contexts Optical Compression research paper by DeepSeek, 2025

DeepSeek · Oct 21, 2025 · Inference and efficiency · 189 citations · 95 upvotes · unverified

Read on arXiv

What it shows

DeepSeek-OCR uses optical 2D mapping to compress long contexts, achieving high OCR precision with reduced vision tokens and demonstrating practical value in document processing.

UnverifiedHugging Face's summary; not yet checked by hand.

More from DeepSeek

All 22
PaperCitations
DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache CompressionA 552B mixture-of-experts model with a 1M-token context, built to shrink the KV cache for agent workloads.Inference and efficiency · Sep 20267
DSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive GenerationDSpark enhances LLM inference speed by combining parallel draft generation with adaptive verification that reduces waste and improves throughput in high-concurrency settings.Inference and efficiency · Jul 2026 · Unverified30
DualPath: Breaking the Storage Bandwidth Bottleneck in Agentic LLM InferenceDualPath addresses KV-cache storage I/O bottlenecks in multi-turn LLM inference by introducing dual-path loading and dynamic load balancing across prefill and decode engines.Agents and evaluation · Feb 2026 · Unverified19
DeepSeek-OCR 2: Visual Causal FlowDeepSeek-OCR 2 introduces DeepEncoder V2 that dynamically reorders visual tokens based on semantic content, enabling more human-like causal reasoning in 2D image understanding through cascaded 1D causal structures.Multimodal and robotics · Jan 2026 · Unverified75
mHC: Manifold-Constrained Hyper-ConnectionsManifold-Constrained Hyper-Connections (mHC) stabilize and scale residual connection architectures by restoring identity mapping properties through manifold projection and infrastructure optimization.Training and scaling · Dec 2025 · Unverified82
DeepSeek-V3.2: Pushing the Frontier of Open Large Language ModelsDeepSeek-V3.2 introduces DeepSeek Sparse Attention and a scalable reinforcement learning framework, achieving superior reasoning and performance compared to GPT-5 and Gemini-3.0-Pro in complex reasoning tasks.Agents and evaluation · Dec 2025 · Unverified761
DeepSeekMath-V2: Towards Self-Verifiable Mathematical ReasoningA self-verifying large language model for theorem proving improves mathematical reasoning by incentivizing rigorous step-by-step derivations and achieves high scores in international competitions.Reasoning · Nov 2025 · Unverified68
Insights into DeepSeek-V3: Scaling Challenges and Reflections on Hardware for AI ArchitecturesDeepSeek-V3 addresses hardware limitations through MLA, MoE, FP8 training, and Multi-Plane Network Topology, enabling efficient large-scale LLM training and inference.Inference and efficiency · May 2025 · Unverified107
Topic
PaperCitations
Efficient Memory Management for Large Language Model Serving with PagedAttentionvLLM manages the KV cache like pages of virtual memory, serving models with 2 to 4 times the throughput.UC Berkeley · Sep 20238,481
FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessFlashAttention computes exact attention with far fewer GPU memory reads and writes, making long sequences faster.Stanford University · May 20225,334
Group Sequence Policy OptimizationGroup Sequence Policy Optimization (GSPO) is a reinforcement learning algorithm that improves training efficiency and performance of large language models by using sequence-level importance ratios and operations.Alibaba (Qwen) · Jul 2025 · Unverified688
SmolVLM: Redefining small and efficient multimodal modelsSmolVLM, a series of compact multimodal models, achieves high performance with minimal GPU memory usage, making efficient deployment on mobile and edge devices possible.Hugging Face · Apr 2025 · Unverified294
Chronos-2: From Univariate to Universal ForecastingChronos-2, a pretrained model with a group attention mechanism, achieves state-of-the-art performance in zero-shot univariate, multivariate, and covariate-informed forecasting tasks.Amazon · Oct 2025 · Unverified178
Fast-dLLM v2: Efficient Block-Diffusion LLMFast-dLLM v2, a block diffusion language model, efficiently converts pretrained autoregressive models for parallel text generation, achieving significant speedup without compromising accuracy.NVIDIA · Sep 2025 · Unverified125
About this paper
Authors
Haoran Wei, Yaofeng Sun, Yukun Li
arXiv
2510.18234 · PDF
Venue
arXiv.org
Citations
189, 20 influential · Semantic Scholar
Upvotes
95 · Hugging Face
Code
github.com/deepseek-ai/DeepSeek-OCR
Lab
DeepSeek · on Companies

Changes

What changed
Influential citationsfirst count: 20Sep 25, 2026
Citationsfirst count: 189Sep 25, 2026
New paperFound by the weekly scan, unverifiedSep 25, 2026

Sources: each lab's own papers and arXiv, with citation and upvote counts from Semantic Scholar and Hugging Face. One-line summaries are for orientation, not a substitute for the paper. Logos via logo.dev; trademarks belong to their owners.

New papers by email

Monday afternoons, only in weeks with new papers from the labs.

Double opt-in. Unsubscribe any time.