HiL-Bench (Human-in-Loop Benchmark): Do Agents Know When to Ask for Help? research paper by Scale AI, 2026
Scale AI · Apr 29, 2026 · Agents and evaluation · 4 upvotes 5 months ago
What it shows
Frontier AI agents struggle with judgment calls about when to seek help, leading to poor performance on incomplete or ambiguous tasks despite having sufficient capabilities.
Hugging Face's summary; not yet checked by hand.
More from Scale AI
All 25Other agents and evaluation papers
TopicAbout this paper
- Authors
- Mohamed Elfeki, Tu Trinh, Kelvin Luu and 9 more
- arXiv
- 2604.09408 · PDF
- Citations
- Not counted yet · Semantic Scholar
- Upvotes
- 4 · Hugging Face
- Code
- github.com/hilbenchauthors/hil-bench
- Lab
- Scale AI · on Companies · on Acquisitions
Changes
| What changed | |
|---|---|
| Sep 26, 2026today | New paperAdded to the listSep 26, 2026today |