Skip to content

Reward Hacking in Rubric-Based Reinforcement Learning research paper by Scale AI, 2026

Scale AI · May 12, 2026 · Reasoning · 4 upvotes 4 months ago

Read on arXiv

What it shows

Research examines reward hacking in rubric-based reinforcement learning, identifying verifier failure and rubric-design limitations as key sources of divergence between training and evaluation metrics.

Hugging Face's summary; not yet checked by hand.

More from Scale AI

All 25
PaperCitations
Root-Cause Attribution Is a Search Problem: Continual Search for Long-Horizon Agent FailuresContinual Search improves automated root-cause attribution in long agent execution traces by iteratively prompting diagnosis until unresolved evidence is found.Agents and evaluation · Sep 20262 weeks ago-
SteerDuplex: Steerable Duplex Speech Dialogue ModelsA full-duplex speech model that follows spoken instructions on tone, persona, pace and voice, with a benchmark of 390 prompts to test it.Multimodal and robotics · Sep 20262 weeks ago-
Studying Without a Syllabus: Task-Agnostic Environment PreprocessingAn agent can explore unfamiliar environments without task-specific guidance to build reusable artifacts that reduce later inference costs, though larger study budgets do not always improve results.Agents and evaluation · Sep 20262 weeks ago-
HarnessOpt-Bench: Evaluating LLMs at Harness OptimizationHarnessOpt-Bench measures how well frontier LLMs iteratively improve agent harnesses under constrained evaluation budgets, revealing substantial variation across models and tasks.Agents and evaluation · Aug 20267 weeks ago-
Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent FailuresAn interaction-centric taxonomy localizes agent failures to specific component interactions to guide targeted repairs across diverse architectures.Agents and evaluation · Jul 20268 weeks ago-
SWE-INTERACT: Reimagining SWE Benchmarks as User-Driven Long-Horizon Coding SessionsSWE-Interact presents a testbed that evaluates coding agents in realistic multi-turn, user-driven software engineering scenarios, revealing significant gaps between single-turn performance and interactive task completion.Agents and evaluation · Jun 20262 months ago-
Not Every Rubric Teaches Equally: Policy-Aware Rubric Rewards for RLVRPOW3R is a policy-aware framework for reinforcement learning with rubric-based rewards that adapts criterion weights during training to improve policy optimization while preserving human-defined criteria importance.Training and scaling · May 20264 months ago-
HiL-Bench (Human-in-Loop Benchmark): Do Agents Know When to Ask for Help?Frontier AI agents struggle with judgment calls about when to seek help, leading to poor performance on incomplete or ambiguous tasks despite having sufficient capabilities.Agents and evaluation · Apr 20265 months ago-
Topic
PaperCitations
Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsAsking a model to write out its intermediate steps (chain of thought) sharply improves its math and logic answers.Google · Jan 20224 years ago22.1k
DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language ModelsDeepSeekMath 7B improves mathematical reasoning through enhanced data pre-training and Group Relative Policy Optimization, achieving high scores on MATH benchmark without external tools.DeepSeek · Feb 2024 · Unverified2 years ago8,994
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement LearningReinforcement learning taught a model long step-by-step reasoning on par with OpenAI o1.DeepSeek · Jan 20251 year ago5,694
DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code IntelligenceDeepSeek-Coder-V2, a Mixture-of-Experts language model, excels in code-specific tasks by enhancing coding and mathematical reasoning capabilities while expanding language support and context length.DeepSeek · Jun 2024 · Unverified2 years ago502
MolmoAct: Action Reasoning Models that can Reason in SpaceAction Reasoning Models (ARMs) integrate perception, planning, and control to enable adaptable and explainable robotic behavior, achieving superior performance across various tasks and settings.Ai2 · Aug 2025 · Unverified1 year ago185
ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent PlanningThinkAct, a dual-system framework, uses reinforced visual latent planning to enable few-shot adaptation, long-horizon planning, and self-correction in embodied AI tasks by bridging high-level reasoning with low-level action execution.NVIDIA · Jul 2025 · Unverified1 year ago172
About this paper
Authors
Anas Mahmoud, MohammadHossein Rezaei, Zihao Wang and 3 more
arXiv
2605.12474 · PDF
Citations
Not counted yet · Semantic Scholar
Upvotes
4 · Hugging Face
Lab
Scale AI · on Companies · on Acquisitions

Changes

What changed
New paperAdded to the listSep 26, 2026today

New papers by email

Monday afternoons, only in weeks with new papers from the labs.

Double opt-in. Unsubscribe any time.