EVOHARNESSBENCH: Can Your Agents Keep Pace with an Evolving Harness? research paper by Salesforce, 2026
Salesforce · Sep 3, 2026 · Agents and evaluation · 32 upvotes 3 weeks ago
What it shows
The study introduces a benchmark to evaluate LLM agents under evolving tool, skill, and agent harnesses, revealing persistent gaps in retention, adaptation, and harness-induced forgetting.
Hugging Face's summary; not yet checked by hand.
More from Salesforce
All 33Other agents and evaluation papers
TopicAbout this paper
- Authors
- Zixuan Ke, Vaidehi Patil, Haizhou Shi and 9 more
- arXiv
- 2609.04280 · PDF
- Citations
- Not counted yet · Semantic Scholar
- Upvotes
- 32 · Hugging Face
- Lab
- Salesforce · on Companies · on Acquisitions · on Quarterly · on Paydays · on TechConf
Changes
| What changed | |
|---|---|
| Sep 26, 2026today | New paperAdded to the listSep 26, 2026today |