LoCoBench-Agent: An Interactive Benchmark for LLM Agents in Long-Context Software Engineering research paper by Salesforce, 2025
Salesforce · Nov 17, 2025 · Agents and evaluation · 3 upvotes 10 months ago
What it shows
LoCoBench-Agent evaluates large language models as autonomous software development agents using interactive scenarios, specialized tools, and multi-turn conversations to assess their long-context performance, comprehension, and efficiency.
Hugging Face's summary; not yet checked by hand.
More from Salesforce
All 33Other agents and evaluation papers
TopicAbout this paper
- Authors
- Jielin Qiu, Zuxin Liu, Zhiwei Liu and 18 more
- arXiv
- 2511.13998 · PDF
- Citations
- Not counted yet · Semantic Scholar
- Upvotes
- 3 · Hugging Face
- Code
- github.com/SalesforceAIResearch/LoCoBench-Agent
- Lab
- Salesforce · on Companies · on Acquisitions · on Quarterly · on Paydays · on TechConf
Changes
| What changed | |
|---|---|
| Sep 26, 2026today | New paperAdded to the listSep 26, 2026today |