Hard2Verify: A Step-Level Verification Benchmark for Open-Ended Frontier Math research paper by Salesforce, 2025
Salesforce · Oct 15, 2025 · Agents and evaluation · 6 upvotes 11 months ago
What it shows
Hard2Verify, a human-annotated benchmark, evaluates step-level verifiers for LLM-based mathematical reasoning systems, highlighting the challenges and performance gaps between open-source and closed-source models.
Hugging Face's summary; not yet checked by hand.
More from Salesforce
All 33Other agents and evaluation papers
TopicAbout this paper
- Authors
- Shrey Pandit, Austin Xu, Xuan-Phi Nguyen and 3 more
- arXiv
- 2510.13744 · PDF
- Citations
- Not counted yet · Semantic Scholar
- Upvotes
- 6 · Hugging Face
- Code
- github.com/SalesforceAIResearch/Hard2Verify
- Lab
- Salesforce · on Companies · on Acquisitions · on Quarterly · on Paydays · on TechConf
Changes
| What changed | |
|---|---|
| Sep 26, 2026today | New paperAdded to the listSep 26, 2026today |