A Careful Examination of Large Language Model Performance on Grade School Arithmetic research paper by Scale AI, 2024
Scale AI · May 1, 2024 · Agents and evaluation · 33 upvotes 2 years ago
What it shows
Evaluation of large language models on a new math benchmark, GSM1k, reveals that many models exhibit signs of overfitting to the existing GSM8k benchmark, leading to performance drops on GSM1k.
Hugging Face's summary; not yet checked by hand.
More from Scale AI
All 25Other agents and evaluation papers
TopicAbout this paper
- Authors
- Hugh Zhang, Jeff Da, Dean Lee and 12 more
- arXiv
- 2405.00332 · PDF
- Citations
- Not counted yet · Semantic Scholar
- Upvotes
- 33 · Hugging Face
- Lab
- Scale AI · on Companies · on Acquisitions
Changes
| What changed | |
|---|---|
| Sep 26, 2026today | New paperAdded to the listSep 26, 2026today |