Skip to content
Papers.

DeepSeek-Prover-V1.5: Harnessing Proof Assistant Feedback for Reinforcement Learning and Monte-Carlo Tree Search research paper by DeepSeek, 2024

DeepSeek · Aug 15, 2024 · Retrieval and data · 207 citations · 63 upvotes · unverified

Read on arXiv

What it shows

DeepSeek-Prover-V1.5 improves theorem proving by optimizing training and inference, utilizing reinforcement learning, and proposing RMaxTS for diverse proof paths, achieving state-of-the-art results on miniF2F and ProofNet benchmarks.

UnverifiedHugging Face's summary; not yet checked by hand.

More from DeepSeek

All 22
PaperCitations
DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache CompressionA 552B mixture-of-experts model with a 1M-token context, built to shrink the KV cache for agent workloads.Inference and efficiency · Sep 20267
DSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive GenerationDSpark enhances LLM inference speed by combining parallel draft generation with adaptive verification that reduces waste and improves throughput in high-concurrency settings.Inference and efficiency · Jul 2026 · Unverified30
DualPath: Breaking the Storage Bandwidth Bottleneck in Agentic LLM InferenceDualPath addresses KV-cache storage I/O bottlenecks in multi-turn LLM inference by introducing dual-path loading and dynamic load balancing across prefill and decode engines.Agents and evaluation · Feb 2026 · Unverified19
DeepSeek-OCR 2: Visual Causal FlowDeepSeek-OCR 2 introduces DeepEncoder V2 that dynamically reorders visual tokens based on semantic content, enabling more human-like causal reasoning in 2D image understanding through cascaded 1D causal structures.Multimodal and robotics · Jan 2026 · Unverified75
mHC: Manifold-Constrained Hyper-ConnectionsManifold-Constrained Hyper-Connections (mHC) stabilize and scale residual connection architectures by restoring identity mapping properties through manifold projection and infrastructure optimization.Training and scaling · Dec 2025 · Unverified82
DeepSeek-V3.2: Pushing the Frontier of Open Large Language ModelsDeepSeek-V3.2 introduces DeepSeek Sparse Attention and a scalable reinforcement learning framework, achieving superior reasoning and performance compared to GPT-5 and Gemini-3.0-Pro in complex reasoning tasks.Agents and evaluation · Dec 2025 · Unverified761
DeepSeekMath-V2: Towards Self-Verifiable Mathematical ReasoningA self-verifying large language model for theorem proving improves mathematical reasoning by incentivizing rigorous step-by-step derivations and achieves high scores in international competitions.Reasoning · Nov 2025 · Unverified68
DeepSeek-OCR: Contexts Optical CompressionDeepSeek-OCR uses optical 2D mapping to compress long contexts, achieving high OCR precision with reduced vision tokens and demonstrating practical value in document processing.Inference and efficiency · Oct 2025 · Unverified189
Topic
PaperCitations
From Local to Global: A Graph RAG Approach to Query-Focused SummarizationGraphRAG builds a knowledge graph of a document set so a model can answer questions about the whole collection.Microsoft · Apr 20242,198
DINOv3DINOv3, a self-supervised learning model, achieves superior performance across various vision tasks by scaling datasets and models, addressing dense feature degradation, and enhancing flexibility with post-hoc strategies.Meta · Aug 2025 · Unverified1,498
Qwen3-VL-Embedding and Qwen3-VL-Reranker: A Unified Framework for State-of-the-Art Multimodal Retrieval and RankingThe Qwen3-VL-Embedding and Qwen3-VL-Reranker models form an end-to-end multimodal search pipeline, leveraging multi-stage training and cross-attention mechanisms to achieve high-precision retrieval across diverse modalities.Alibaba (Qwen) · Jan 2026 · Unverified231
Learning to Discover at Test TimeTest-time training enables AI systems to discover optimal solutions for specific scientific problems through continual learning focused on individual challenges rather than generalization.Stanford University · Jan 2026 · Unverified82
Arctic-Embed: Scalable, Efficient, and Accurate Text Embedding ModelsOpen text embedding models from 22M to 334M parameters that led MTEB retrieval for their size at release.Snowflake · May 202480
Pico-Banana-400K: A Large-Scale Dataset for Text-Guided Image EditingPico-Banana-400K is a large-scale, high-quality dataset for instruction-based image editing, featuring diverse edit pairs, multi-turn editing, preference subsets, and long-short instruction pairs, enabling comprehensive research and benchmarking.Apple · Oct 2025 · Unverified62
About this paper
Authors
Huajian Xin, Z. Z. Ren, Junxiao Song and 14 more
arXiv
2408.08152 · PDF
Venue
arXiv.org
Citations
207, 28 influential · Semantic Scholar
Upvotes
63 · Hugging Face
Code
github.com/deepseek-ai/deepseek-prover-v1.5
Lab
DeepSeek · on Companies

Changes

What changed
Influential citationsfirst count: 28Sep 25, 2026
Citationsfirst count: 207Sep 25, 2026
New paperFound by the weekly scan, unverifiedSep 25, 2026

Sources: each lab's own papers and arXiv, with citation and upvote counts from Semantic Scholar and Hugging Face. One-line summaries are for orientation, not a substitute for the paper. Logos via logo.dev; trademarks belong to their owners.

New papers by email

Monday afternoons, only in weeks with new papers from the labs.

Double opt-in. Unsubscribe any time.