Skip to content

Planning In Natural Language Improves LLM Search For Code Generation research paper by Scale AI, 2024

Scale AI · Sep 5, 2024 · Retrieval and data 2 years ago

Read on arXiv

What it shows

PLANSEARCH, a novel search algorithm, improves performance in coding benchmarks by generating a diverse set of natural language plans, outperforming existing methods through increased diversity in solutions.

Hugging Face's summary; not yet checked by hand.

More from Scale AI

All 25
PaperCitations
Root-Cause Attribution Is a Search Problem: Continual Search for Long-Horizon Agent FailuresContinual Search improves automated root-cause attribution in long agent execution traces by iteratively prompting diagnosis until unresolved evidence is found.Agents and evaluation · Sep 20262 weeks ago-
SteerDuplex: Steerable Duplex Speech Dialogue ModelsA full-duplex speech model that follows spoken instructions on tone, persona, pace and voice, with a benchmark of 390 prompts to test it.Multimodal and robotics · Sep 20262 weeks ago-
Studying Without a Syllabus: Task-Agnostic Environment PreprocessingAn agent can explore unfamiliar environments without task-specific guidance to build reusable artifacts that reduce later inference costs, though larger study budgets do not always improve results.Agents and evaluation · Sep 20262 weeks ago-
HarnessOpt-Bench: Evaluating LLMs at Harness OptimizationHarnessOpt-Bench measures how well frontier LLMs iteratively improve agent harnesses under constrained evaluation budgets, revealing substantial variation across models and tasks.Agents and evaluation · Aug 20267 weeks ago-
Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent FailuresAn interaction-centric taxonomy localizes agent failures to specific component interactions to guide targeted repairs across diverse architectures.Agents and evaluation · Jul 20268 weeks ago-
SWE-INTERACT: Reimagining SWE Benchmarks as User-Driven Long-Horizon Coding SessionsSWE-Interact presents a testbed that evaluates coding agents in realistic multi-turn, user-driven software engineering scenarios, revealing significant gaps between single-turn performance and interactive task completion.Agents and evaluation · Jun 20262 months ago-
Not Every Rubric Teaches Equally: Policy-Aware Rubric Rewards for RLVRPOW3R is a policy-aware framework for reinforcement learning with rubric-based rewards that adapts criterion weights during training to improve policy optimization while preserving human-defined criteria importance.Training and scaling · May 20264 months ago-
Reward Hacking in Rubric-Based Reinforcement LearningResearch examines reward hacking in rubric-based reinforcement learning, identifying verifier failure and rubric-design limitations as key sources of divergence between training and evaluation metrics.Reasoning · May 20264 months ago-
Topic
PaperCitations
From Local to Global: A Graph RAG Approach to Query-Focused SummarizationGraphRAG builds a knowledge graph of a document set so a model can answer questions about the whole collection.Microsoft · Apr 20242 years ago2,198
DINOv3DINOv3, a self-supervised learning model, achieves superior performance across various vision tasks by scaling datasets and models, addressing dense feature degradation, and enhancing flexibility with post-hoc strategies.Meta · Aug 2025 · Unverified1 year ago1,498
DeepSeek-Prover: Advancing Theorem Proving in LLMs through Large-Scale Synthetic DataGenerating extensive Lean 4 proof data from mathematical competition problems improved DeepSeekMath 7B's theorem-proving capabilities over GPT-4 and other methods.DeepSeek · May 2024 · Unverified2 years ago260
Qwen3-VL-Embedding and Qwen3-VL-Reranker: A Unified Framework for State-of-the-Art Multimodal Retrieval and RankingThe Qwen3-VL-Embedding and Qwen3-VL-Reranker models form an end-to-end multimodal search pipeline, leveraging multi-stage training and cross-attention mechanisms to achieve high-precision retrieval across diverse modalities.Alibaba (Qwen) · Jan 2026 · Unverified8 months ago231
DeepSeek-Prover-V1.5: Harnessing Proof Assistant Feedback for Reinforcement Learning and Monte-Carlo Tree SearchDeepSeek-Prover-V1.5 improves theorem proving by optimizing training and inference, utilizing reinforcement learning, and proposing RMaxTS for diverse proof paths, achieving state-of-the-art results on miniF2F and ProofNet benchmarks.DeepSeek · Aug 2024 · Unverified2 years ago207
Learning to Discover at Test TimeTest-time training enables AI systems to discover optimal solutions for specific scientific problems through continual learning focused on individual challenges rather than generalization.Stanford University · Jan 2026 · Unverified8 months ago82
About this paper
Authors
Evan Wang, Federico Cassano, Catherine Wu and 7 more
arXiv
2409.03733 · PDF
Citations
Not counted yet · Semantic Scholar
Upvotes
0 · Hugging Face
Lab
Scale AI · on Companies · on Acquisitions

Changes

What changed
New paperAdded to the listSep 26, 2026today

New papers by email

Monday afternoons, only in weeks with new papers from the labs.

Double opt-in. Unsubscribe any time.