Skip to content

SciPredict: Can LLMs Predict the Outcomes of Scientific Experiments in Natural Sciences? research paper by Scale AI, 2026

Scale AI · Apr 12, 2026 · Applied AI · 4 upvotes 5 months ago

Read on arXiv

What it shows

SciPredict benchmark reveals that large language models struggle to accurately predict scientific experiment outcomes and cannot reliably assess prediction confidence, unlike human experts who show better calibration and performance when experiments are deemed predictable.

Hugging Face's summary; not yet checked by hand.

More from Scale AI

All 25
PaperCitations
Root-Cause Attribution Is a Search Problem: Continual Search for Long-Horizon Agent FailuresContinual Search improves automated root-cause attribution in long agent execution traces by iteratively prompting diagnosis until unresolved evidence is found.Agents and evaluation · Sep 20262 weeks ago-
SteerDuplex: Steerable Duplex Speech Dialogue ModelsA full-duplex speech model that follows spoken instructions on tone, persona, pace and voice, with a benchmark of 390 prompts to test it.Multimodal and robotics · Sep 20262 weeks ago-
Studying Without a Syllabus: Task-Agnostic Environment PreprocessingAn agent can explore unfamiliar environments without task-specific guidance to build reusable artifacts that reduce later inference costs, though larger study budgets do not always improve results.Agents and evaluation · Sep 20262 weeks ago-
HarnessOpt-Bench: Evaluating LLMs at Harness OptimizationHarnessOpt-Bench measures how well frontier LLMs iteratively improve agent harnesses under constrained evaluation budgets, revealing substantial variation across models and tasks.Agents and evaluation · Aug 20267 weeks ago-
Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent FailuresAn interaction-centric taxonomy localizes agent failures to specific component interactions to guide targeted repairs across diverse architectures.Agents and evaluation · Jul 20268 weeks ago-
SWE-INTERACT: Reimagining SWE Benchmarks as User-Driven Long-Horizon Coding SessionsSWE-Interact presents a testbed that evaluates coding agents in realistic multi-turn, user-driven software engineering scenarios, revealing significant gaps between single-turn performance and interactive task completion.Agents and evaluation · Jun 20262 months ago-
Not Every Rubric Teaches Equally: Policy-Aware Rubric Rewards for RLVRPOW3R is a policy-aware framework for reinforcement learning with rubric-based rewards that adapts criterion weights during training to improve policy optimization while preserving human-defined criteria importance.Training and scaling · May 20264 months ago-
Reward Hacking in Rubric-Based Reinforcement LearningResearch examines reward hacking in rubric-based reinforcement learning, identifying verifier failure and rubric-design limitations as key sources of divergence between training and evaluation metrics.Reasoning · May 20264 months ago-
Topic
About this paper
Authors
Udari Madhushani Sehwag, Elaine Lau, Haniyeh Ehsani Oskouie and 14 more
arXiv
2604.10718 · PDF
Citations
Not counted yet · Semantic Scholar
Upvotes
4 · Hugging Face
Lab
Scale AI · on Companies · on Acquisitions

Changes

What changed
New paperAdded to the listSep 26, 2026today

New papers by email

Monday afternoons, only in weeks with new papers from the labs.

Double opt-in. Unsubscribe any time.