Skip to content

ActiveSaddler: Automated Curriculum Learning for Agent Harness Optimization research paper by Microsoft, 2026

Microsoft · Oct 1, 2026 · Agents and evaluation · 66 upvotes · unverified 4 days ago

Read on arXiv

What it shows

Automated harness optimization can substantially improve LLM agents by iteratively updating their prompts, tool interfaces, and control logic from execution feedback.

By Sungho Park, Wonjoong Kim, Jue Zhang and 8 more · arXiv 2610.00906 · PDF

UnverifiedHugging Face's summary; not yet checked by hand.

More from Microsoft

All 49
PaperCitations
Follow the Entities: A Corpus Map for Agentic SearchAnswering questions and completing tasks over large document collections often requires connecting evidence spread across multiple documents, such as a project's approval recorded in one, its requirements in another,...Agents and evaluation · Sep 2026 · Unverified6 days ago-
Agensh: Scaling Organizational Intelligence to 1,024 AgentsA multi-agent system can reduce latency on complex tasks by executing work concurrently.Agents and evaluation · Sep 2026 · Unverified13 days ago0
The Tasteful Agent: Measuring and Improving Taste in Long-Horizon TasksLLM agents increasingly work on long-horizon tasks, and the decisions they make along the way, such as which hypothesis to test or which implementation to build on, determine the outcome of the whole run.Agents and evaluation · Sep 2026 · Unverified13 days ago1
When EOS Tokens Disagree: Understanding Length Inflation in On-Policy DistillationWe study length inflation in on-policy distillation (OPD), where student responses can become excessively long and even exhaust the generation budget.Foundation models · Sep 2026 · Unverified2 weeks ago2
When2Think: Learning Difficulty-Aware Length Control for Efficient Hybrid Reasoning ModelsLarge Reasoning Models (LRMs) achieve strong performance on complex tasks but exhibit systematic inefficiency: they often overthink easy problems and underthink hard ones.Reasoning · Sep 2026 · Unverified2 weeks ago0
BI-Agent and BI-Bench: Towards Automating End-to-End Business IntelligenceBusiness intelligence (BI) is a cornerstone of enterprise decision-making and is widely used by enterprise users in software such as Power BI and Tableau.Agents and evaluation · Sep 2026 · Unverified2 weeks ago0
StudentSim: Training LLM-based Student SimulatorsStudentSim trains per-student simulators that answer like a given learner and change their answers under a tutor's guidance.Applied AI · Sep 20264 weeks ago0
AutoSaddler: Automatic Harness Optimization with Durable Updates from Agent Execution TracesAutoSaddler automatically improves LLM agent harnesses via offline failure-driven optimization, boosting performance on long-horizon benchmarks.Agents and evaluation · Aug 2026 · Unverified6 weeks ago11
Topic
PaperCitations
Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic CapabilitiesGemini 2.X model family, including Gemini 2.5 Pro and Flash, offers superior coding, reasoning, and multimodal understanding capabilities across a range of computational efficiencies.Google · Jul 2025 · Unverified1 year ago4,244
DeepSeek-V3.2: Pushing the Frontier of Open Large Language ModelsDeepSeek-V3.2 introduces DeepSeek Sparse Attention and a scalable reinforcement learning framework, achieving superior reasoning and performance compared to GPT-5 and Gemini-3.0-Pro in complex reasoning tasks.DeepSeek · Dec 2025 · Unverified10 months ago784
Aya Model: An Instruction Finetuned Open-Access Multilingual Language ModelAya, a multilingual generative language model supporting over 50% lower-resourced languages, excels in both generative and discriminative tasks across 99 languages, and provides extensive evaluations and open-source resources.Cohere · Feb 20242 years ago410
WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?WorkArena and BrowserGym evaluate large language model-based agents' ability to perform enterprise software tasks, revealing gaps in current agent capabilities and differences between open and closed-source LLMs.ServiceNow · Mar 20242 years ago387
Rubrics as Rewards: Reinforcement Learning Beyond Verifiable DomainsRubrics as Rewards (RaR) framework uses structured rubrics as interpretable reward signals for on-policy training, improving performance in real-world reinforcement learning tasks with subjective criteria.Scale AI · Jul 20251 year ago343
SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?SWE-Bench Pro is a challenging benchmark for coding models, featuring complex, enterprise-level problems that require substantial code modifications, with performance evaluations showing significant limitations in current models.Scale AI · Sep 20251 year ago270
About this paper
Authors
Sungho Park, Wonjoong Kim, Jue Zhang and 8 more
arXiv
2610.00906 · PDF
Citations
Not counted yet · Semantic Scholar
Upvotes
66 · Hugging Face
Lab
Microsoft · on Companies · on Acquisitions · on Quarterly · on Paydays · on Releases · on TechConf

Changes

What changed
New paperFound by the weekly scan, unverifiedOct 5, 2026today

New papers by email

Monday afternoons, only in weeks with new papers from the labs.

Double opt-in. Unsubscribe any time.