WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks? research paper by ServiceNow, 2024
ServiceNow · Mar 12, 2024 · Agents and evaluation · 2 upvotes 2 years ago
What it shows
WorkArena and BrowserGym evaluate large language model-based agents' ability to perform enterprise software tasks, revealing gaps in current agent capabilities and differences between open and closed-source LLMs.
Hugging Face's summary; not yet checked by hand.
More from ServiceNow
All 20Other agents and evaluation papers
TopicAbout this paper
- Authors
- Alexandre Drouin, Maxime Gasse, Massimo Caccia and 9 more
- arXiv
- 2403.07718 · PDF
- Citations
- Not counted yet · Semantic Scholar
- Upvotes
- 2 · Hugging Face
- Lab
- ServiceNow · on Companies · on Acquisitions · on Quarterly · on Paydays · on TechConf
Changes
| What changed | |
|---|---|
| Sep 26, 2026today | New paperAdded to the listSep 26, 2026today |