MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers research paper by Scale AI, 2026
Scale AI · Jan 31, 2026 · Agents and evaluation · 1 upvotes 7 months ago
What it shows
MCP-Atlas is a large-scale benchmark for evaluating tool-use competency in LLMs, featuring 36 real servers and 220 tools across 1,000 realistic multi-step tasks with claims-based scoring.
Hugging Face's summary; not yet checked by hand.
More from Scale AI
All 25Other agents and evaluation papers
TopicAbout this paper
- Authors
- Chaithanya Bandi, Ben Hertzberg, Geobio Boo and 13 more
- arXiv
- 2602.00933 · PDF
- Citations
- Not counted yet · Semantic Scholar
- Upvotes
- 1 · Hugging Face
- Lab
- Scale AI · on Companies · on Acquisitions
Changes
| What changed | |
|---|---|
| Sep 26, 2026today | New paperAdded to the listSep 26, 2026today |