The Tasteful Agent: Measuring and Improving Taste in Long-Horizon Tasks research paper by Microsoft, 2026
Microsoft · Sep 22, 2026 · Agents and evaluation · 0 citations · 119 upvotes · unverified
What it shows
LLM agents increasingly work on long-horizon tasks, and the decisions they make along the way, such as which hypothesis to test or which implementation to build on, determine the outcome of the whole run.
UnverifiedHugging Face's summary; not yet checked by hand.
More from Microsoft
All 45Other agents and evaluation papers
TopicAbout this paper
- Authors
- Wenbo Pan, Zhichao Liu, Shujie Liu and 6 more
- arXiv
- 2609.25804 · PDF
- Citations
- 0, 0 influential · Semantic Scholar
- Upvotes
- 119 · Hugging Face
- Code
- github.com/wbopan/tastebench
- Lab
- Microsoft · on Companies · on Acquisitions · on Quarterly · on Paydays · on Releases · on TechConf
Changes
| What changed | |
|---|---|
| Sep 25, 2026 | Upvotes118 to 119 (+1)Sep 25, 2026 |
| Sep 25, 2026 | Influential citationsfirst count: 0Sep 25, 2026 |
| Sep 25, 2026 | Citationsfirst count: 0Sep 25, 2026 |
| Sep 25, 2026 | New paperFound by the weekly scan, unverifiedSep 25, 2026 |
Sources: each lab's own papers and arXiv, with citation and upvote counts from Semantic Scholar and Hugging Face. One-line summaries are for orientation, not a substitute for the paper. Logos via logo.dev; trademarks belong to their owners.