Skip to content

Selecting Diverse SFT Traces Improves Post-RL Generalization research paper by Google, 2026

Google · Sep 27, 2026 · Reasoning · 38 upvotes · unverified 8 days ago

Read on arXiv

What it shows

Verified solutions are not equally useful for preparing reasoning models for reinforcement learning (RL).

By Dylan Zhang, Mingyuan Wu, Jinning Li · arXiv 2609.33780 · PDF

UnverifiedHugging Face's summary; not yet checked by hand.

More from Google

All 36
PaperCitations
RPTune: Learned Context Curation for LLM Catalog SearchFor small merchant businesses (SMBs) whose catalogs fit within a long-context LLM, full-catalog prompting offers a compelling alternative to multi-stage retrieval designed primarily for large marketplaces with...Retrieval and data · Oct 2026 · Unverified4 days ago-
AIM: Agentic Idea Management for Automated ResearchFrontier LLMs are increasingly used to automate scientific research through iterative search.Agents and evaluation · Sep 2026 · Unverified6 days ago-
RRSI: Regularized Recursive Self-Improvement of Agent HarnessesRRSI keeps self-improving agent harnesses from memorising their training tasks, so gains carry over to new benchmarks.Agents and evaluation · Sep 20262 weeks ago3
Verifiable Social Reasoning for LLM AssistantsLLM assistants are widely used for daily social advice, yet evaluating their social reasoning in such consultation settings remains challenging since (i) it requires setups where the assistant learns about social...Reasoning · Sep 2026 · Unverified2 weeks ago1
Dream-RSI: Recursive Self-Improvement through Evolving WorldsDream-RSI enables scalable recursive self-improvement by using historical discovery replay to evaluate exploration policies offline, reducing costly online evaluations.Retrieval and data · Sep 2026 · Unverified3 weeks ago9
Procedural Graphs: Self-Evolving Execution Structures for LLM AgentsA procedural graph framework organizes agent actions into structured relational triplets, providing situational guidance and self-evolving topology to improve long-horizon tool use.Foundation models · Sep 2026 · Unverified3 weeks ago0
WikiSkill: Compiling Agent Experience into Persistent Knowledge for Skill EvolutionWikiSkill co-evolves reusable agent skills with a persistent knowledge base to systematically accumulate experience and improve performance across models.Agents and evaluation · Aug 2026 · Unverified5 weeks ago6
EnvHarness: Awakening Static Worlds for Agent LearningEnvHarness and EnvRigger dynamically reshape static environments via programmable plugins to target agent weaknesses and improve reinforcement learning co-evolution.Agents and evaluation · Aug 2026 · Unverified6 weeks ago8
Topic
PaperCitations
DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language ModelsDeepSeekMath 7B improves mathematical reasoning through enhanced data pre-training and Group Relative Policy Optimization, achieving high scores on MATH benchmark without external tools.DeepSeek · Feb 2024 · Unverified2 years ago9,498
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement LearningReinforcement learning taught a model long step-by-step reasoning on par with OpenAI o1.DeepSeek · Jan 20251 year ago5,805
DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code IntelligenceDeepSeek-Coder-V2, a Mixture-of-Experts language model, excels in code-specific tasks by enhancing coding and mathematical reasoning capabilities while expanding language support and context length.DeepSeek · Jun 2024 · Unverified2 years ago507
MolmoAct: Action Reasoning Models that can Reason in SpaceAction Reasoning Models (ARMs) integrate perception, planning, and control to enable adaptable and explainable robotic behavior, achieving superior performance across various tasks and settings.Ai2 · Aug 2025 · Unverified1 year ago199
ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent PlanningThinkAct, a dual-system framework, uses reinforced visual latent planning to enable few-shot adaptation, long-horizon planning, and self-correction in embodied AI tasks by bridging high-level reasoning with low-level action execution.NVIDIA · Jul 2025 · Unverified1 year ago177
WebShaper: Agentically Data Synthesizing via Information-Seeking FormalizationWebShaper, a formalization-driven framework, synthesizes information-seeking datasets using set theory and Knowledge Projections to enhance reasoning structure and achieve top performance in open-sourced benchmarks.Alibaba (Qwen) · Jul 2025 · Unverified1 year ago115
About this paper
Authors
Dylan Zhang, Mingyuan Wu, Jinning Li
arXiv
2609.33780 · PDF
Citations
Not counted yet · Semantic Scholar
Upvotes
38 · Hugging Face
Lab
Google · on Companies · on Acquisitions · on Paydays · on TechConf · on Releases

Changes

What changed
New paperFound by the weekly scan, unverifiedOct 5, 2026today

New papers by email

Monday afternoons, only in weeks with new papers from the labs.

Double opt-in. Unsubscribe any time.