Skip to content

Scale AI research papers

Company · 25 papers · 0 citations · latest Sep 2026 2 weeks ago

Research page

25 of 25 papers, newest first

PaperCitations
Root-Cause Attribution Is a Search Problem: Continual Search for Long-Horizon Agent FailuresContinual Search improves automated root-cause attribution in long agent execution traces by iteratively prompting diagnosis until unresolved evidence is found.Agents and evaluation · Sep 20262 weeks ago-
SteerDuplex: Steerable Duplex Speech Dialogue ModelsA full-duplex speech model that follows spoken instructions on tone, persona, pace and voice, with a benchmark of 390 prompts to test it.Multimodal and robotics · Sep 20262 weeks ago-
Studying Without a Syllabus: Task-Agnostic Environment PreprocessingAn agent can explore unfamiliar environments without task-specific guidance to build reusable artifacts that reduce later inference costs, though larger study budgets do not always improve results.Agents and evaluation · Sep 20262 weeks ago-
HarnessOpt-Bench: Evaluating LLMs at Harness OptimizationHarnessOpt-Bench measures how well frontier LLMs iteratively improve agent harnesses under constrained evaluation budgets, revealing substantial variation across models and tasks.Agents and evaluation · Aug 20267 weeks ago-
Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent FailuresAn interaction-centric taxonomy localizes agent failures to specific component interactions to guide targeted repairs across diverse architectures.Agents and evaluation · Jul 20268 weeks ago-
SWE-INTERACT: Reimagining SWE Benchmarks as User-Driven Long-Horizon Coding SessionsSWE-Interact presents a testbed that evaluates coding agents in realistic multi-turn, user-driven software engineering scenarios, revealing significant gaps between single-turn performance and interactive task completion.Agents and evaluation · Jun 20262 months ago-
Not Every Rubric Teaches Equally: Policy-Aware Rubric Rewards for RLVRPOW3R is a policy-aware framework for reinforcement learning with rubric-based rewards that adapts criterion weights during training to improve policy optimization while preserving human-defined criteria importance.Training and scaling · May 20264 months ago-
Reward Hacking in Rubric-Based Reinforcement LearningResearch examines reward hacking in rubric-based reinforcement learning, identifying verifier failure and rubric-design limitations as key sources of divergence between training and evaluation metrics.Reasoning · May 20264 months ago-
HiL-Bench (Human-in-Loop Benchmark): Do Agents Know When to Ask for Help?Frontier AI agents struggle with judgment calls about when to seek help, leading to poor performance on incomplete or ambiguous tasks despite having sufficient capabilities.Agents and evaluation · Apr 20265 months ago-
SciPredict: Can LLMs Predict the Outcomes of Scientific Experiments in Natural Sciences?SciPredict benchmark reveals that large language models struggle to accurately predict scientific experiment outcomes and cannot reliably assess prediction confidence, unlike human experts who show better calibration and performance when experiments are deemed predictable.Applied AI · Apr 20265 months ago-
MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP ServersMCP-Atlas is a large-scale benchmark for evaluating tool-use competency in LLMs, featuring 36 real servers and 220 tools across 1,000 realistic multi-step tasks with claims-based scoring.Agents and evaluation · Jan 20267 months ago-
Agentic Rubrics as Contextual Verifiers for SWE AgentsAgentic Rubrics enable efficient and scalable verification for software engineering agents by creating context-aware checklists that outperform traditional methods while maintaining interpretability.Agents and evaluation · Jan 20268 months ago-
PropensityBench: Evaluating Latent Safety Risks in Large Language Models via an Agentic ApproachA new benchmark framework, PropensityBench, evaluates models' likelihood to engage in risky behaviors when equipped with simulated dangerous capabilities, highlighting the need for dynamic propensity assessments in AI safety.Agents and evaluation · Nov 202510 months ago-
ResearchRubrics: A Benchmark of Prompts and Rubrics For Evaluating Deep Research AgentsResearchRubrics is a benchmark for evaluating deep research agents, using expert rubrics to assess their factual grounding, reasoning, and clarity across diverse, complex tasks.Agents and evaluation · Nov 202510 months ago-
MoReBench: Evaluating Procedural and Pluralistic Moral Reasoning in Language Models, More than OutcomesMoReBench and MoReBench-Theory provide benchmarks for evaluating AI's moral reasoning and decision-making processes, highlighting the need for process-focused evaluation and transparency in AI systems.Reasoning · Oct 202511 months ago-
TutorBench: A Benchmark To Assess Tutoring Capabilities Of Large Language ModelsTutorBench is a dataset and benchmark for evaluating the tutoring skills of large language models, showing significant room for improvement in adaptive explanations, feedback, and active learning.Agents and evaluation · Oct 202511 months ago-
Chasing the Tail: Effective Rubric-based Reward Modeling for Large Language Model Post-TrainingRubric-based rewards mitigate reward over-optimization in reinforcement fine-tuning by leveraging off-policy examples while maintaining reward reliability.Training and scaling · Sep 20251 year ago-
SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?SWE-Bench Pro is a challenging benchmark for coding models, featuring complex, enterprise-level problems that require substantial code modifications, with performance evaluations showing significant limitations in current models.Agents and evaluation · Sep 20251 year ago-
Rubrics as Rewards: Reinforcement Learning Beyond Verifiable DomainsRubrics as Rewards (RaR) framework uses structured rubrics as interpretable reward signals for on-policy training, improving performance in real-world reinforcement learning tasks with subjective criteria.Agents and evaluation · Jul 20251 year ago-
EnigmaEval: A Benchmark of Long Multimodal Reasoning ChallengesEnigmaEval introduces a benchmark of complex multimodal puzzles to evaluate language models on implicit knowledge synthesis and multi-step deductive reasoning.Agents and evaluation · Feb 20251 year ago-
MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMsMultiChallenge evaluates large language models on multi-turn conversations, identifying four key challenges that current models struggle with, despite high performance on existing benchmarks.Agents and evaluation · Jan 20251 year ago-
Refusal-Trained LLMs Are Easily Jailbroken As Browser AgentsBrowserART, a test suite for red-teaming browser agents, reveals that LLMs trained to refuse harmful instructions in chats often fail to do so when equipped with web browser capabilities.Agents and evaluation · Oct 20241 year ago-
Planning In Natural Language Improves LLM Search For Code GenerationPLANSEARCH, a novel search algorithm, improves performance in coding benchmarks by generating a diverse set of natural language plans, outperforming existing methods through increased diversity in solutions.Retrieval and data · Sep 20242 years ago-
LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks YetMulti-turn human jailbreaks reveal significant vulnerabilities in large language model defenses, achieving high attack success rates and exposing flaws in machine unlearning mechanisms.Foundation models · Aug 20242 years ago-
A Careful Examination of Large Language Model Performance on Grade School ArithmeticEvaluation of large language models on a new math benchmark, GSM1k, reveals that many models exhibit signs of overfitting to the existing GSM8k benchmark, leading to performance drops on GSM1k.Agents and evaluation · May 20242 years ago-
About and links

New papers by email

Monday afternoons, only in weeks with new papers from the labs.

Double opt-in. Unsubscribe any time.