HarnessOpt-Bench: Evaluating LLMs at Harness Optimization research paper by Scale AI, 2026
Scale AI · Aug 6, 2026 · Agents and evaluation · 35 upvotes 7 weeks ago
What it shows
HarnessOpt-Bench measures how well frontier LLMs iteratively improve agent harnesses under constrained evaluation budgets, revealing substantial variation across models and tasks.
Hugging Face's summary; not yet checked by hand.
More from Scale AI
All 25Other agents and evaluation papers
TopicAbout this paper
- Authors
- Varun Ursekar, Apaar Shanker, Yash Maurya and 4 more
- arXiv
- 2608.06301 · PDF
- Citations
- Not counted yet · Semantic Scholar
- Upvotes
- 35 · Hugging Face
- Lab
- Scale AI · on Companies · on Acquisitions
Changes
| What changed | |
|---|---|
| Sep 26, 2026today | New paperAdded to the listSep 26, 2026today |