SWE-INTERACT: Reimagining SWE Benchmarks as User-Driven Long-Horizon Coding Sessions research paper by Scale AI, 2026
Scale AI · Jun 29, 2026 · Agents and evaluation · 7 upvotes 2 months ago
What it shows
SWE-Interact presents a testbed that evaluates coding agents in realistic multi-turn, user-driven software engineering scenarios, revealing significant gaps between single-turn performance and interactive task completion.
Hugging Face's summary; not yet checked by hand.
More from Scale AI
All 25Other agents and evaluation papers
TopicAbout this paper
- Authors
- Mohit Raghavendra, Anisha Gunjal, Aakash Sabharwal and 1 more
- arXiv
- 2606.30573 · PDF
- Citations
- Not counted yet · Semantic Scholar
- Upvotes
- 7 · Hugging Face
- Code
- github.com/scaleapi/SWE-Interact
- Lab
- Scale AI · on Companies · on Acquisitions
Changes
| What changed | |
|---|---|
| Sep 26, 2026today | New paperAdded to the listSep 26, 2026today |