SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks? research paper by Scale AI, 2025
Scale AI · Sep 21, 2025 · Agents and evaluation · 21 upvotes 1 year ago
What it shows
SWE-Bench Pro is a challenging benchmark for coding models, featuring complex, enterprise-level problems that require substantial code modifications, with performance evaluations showing significant limitations in current models.
Hugging Face's summary; not yet checked by hand.
More from Scale AI
All 25Other agents and evaluation papers
TopicAbout this paper
- Authors
- Xiang Deng, Jeff Da, Edwin Pan and 16 more
- arXiv
- 2509.16941 · PDF
- Citations
- Not counted yet · Semantic Scholar
- Upvotes
- 21 · Hugging Face
- Code
- github.com/scaleapi/SWE-bench_Pro-os
- Lab
- Scale AI · on Companies · on Acquisitions
Changes
| What changed | |
|---|---|
| Sep 26, 2026today | New paperAdded to the listSep 26, 2026today |