Skip to content

Mid-Harness: Scaling Actions Between Model and Harness for Terminal Agents research paper by NVIDIA, 2026

NVIDIA · Sep 30, 2026 · Agents and evaluation · 112 upvotes · unverified 5 days ago

Read on arXiv

What it shows

Terminal agents act through stochastic model generations, yet the ability to generate a useful action does not ensure its reliable execution.

By Minki Kang, Ryo Hachiuma, Shaokun Zhang and 8 more · arXiv 2609.39982 · PDF

UnverifiedHugging Face's summary; not yet checked by hand.

More from NVIDIA

All 79
PaperCitations
LongLive-Plug: Once-for-All Distillation for Video GenerationVideo diffusion models are increasingly developed into specialized models for diverse downstream tasks, and this development often includes a distillation stage, for example to accelerate sampling or to improve...Multimodal and robotics · Sep 2026 · Unverified6 days ago-
PixelUMM: Encoder-Free Unified Image and Video Understanding and GenerationUnified Multimodal Models (UMMs) often rely on separate visual representations for understanding and generation, increasing visual context length and complicating integration with established vision-language...Multimodal and robotics · Sep 2026 · Unverified6 days ago-
An RL View of OPD: Least Square Policy Distillation for Sample-Efficient LLM ReasoningWe study on-policy distillation (OPD) through the lens of reinforcement learning, establishing a connection between the reverse-KL objective in OPD and KL-regularized policy optimization.Reasoning · Sep 2026 · Unverified7 days ago-
SoL-Pi: Recursively Scaling Auto-Research Loops for Efficient Agent HarnessAs coding agents move from supervised code completion to unattended, around-the-clock exploration, their work expands from isolated predictions into long trajectories of reasoning, tool use, and feedback.Agents and evaluation · Sep 2026 · Unverified2 weeks ago4
An Open Recipe for IMO Gold: Training Nemotron for Olympiad MathematicsA natural-language proof-generation pipeline using post-trained Nemotron 3 Ultra checkpoints achieves gold-medal performance on IMO 2026 through iterative verification and refinement without external tools.Training and scaling · Sep 2026 · Unverified3 weeks ago0
Online Draft Co-Training for Speculative Decoding in Large-Scale, Long-Context RL Post-TrainingA system for large-scale online draft co-training accelerates speculative decoding in RL post-training by extending context-parallel attention and adding cross-stage feature transport.Training and scaling · Sep 2026 · Unverified4 weeks ago2
Sol-Attn: Accelerating Video Generation Inference via On-the-Fly Attention SparsificationSol-Attn improves training-free sparse attention for diffusion transformers by combining dynamic block routing, sparse computation, and approximation correction in a single online pass to accelerate video generation without sacrificing quality.Inference and efficiency · Jul 2026 · Unverified2 months ago6
SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video GenerationSANA-Video 2.0 is a hybrid video diffusion transformer that combines linear and softmax attention to generate high-resolution video efficiently on a single GPU.Inference and efficiency · Jul 2026 · Unverified2 months ago2
Topic
PaperCitations
Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic CapabilitiesGemini 2.X model family, including Gemini 2.5 Pro and Flash, offers superior coding, reasoning, and multimodal understanding capabilities across a range of computational efficiencies.Google · Jul 2025 · Unverified1 year ago4,244
DeepSeek-V3.2: Pushing the Frontier of Open Large Language ModelsDeepSeek-V3.2 introduces DeepSeek Sparse Attention and a scalable reinforcement learning framework, achieving superior reasoning and performance compared to GPT-5 and Gemini-3.0-Pro in complex reasoning tasks.DeepSeek · Dec 2025 · Unverified10 months ago784
Aya Model: An Instruction Finetuned Open-Access Multilingual Language ModelAya, a multilingual generative language model supporting over 50% lower-resourced languages, excels in both generative and discriminative tasks across 99 languages, and provides extensive evaluations and open-source resources.Cohere · Feb 20242 years ago410
WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?WorkArena and BrowserGym evaluate large language model-based agents' ability to perform enterprise software tasks, revealing gaps in current agent capabilities and differences between open and closed-source LLMs.ServiceNow · Mar 20242 years ago387
Rubrics as Rewards: Reinforcement Learning Beyond Verifiable DomainsRubrics as Rewards (RaR) framework uses structured rubrics as interpretable reward signals for on-policy training, improving performance in real-world reinforcement learning tasks with subjective criteria.Scale AI · Jul 20251 year ago343
SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?SWE-Bench Pro is a challenging benchmark for coding models, featuring complex, enterprise-level problems that require substantial code modifications, with performance evaluations showing significant limitations in current models.Scale AI · Sep 20251 year ago270
About this paper
Authors
Minki Kang, Ryo Hachiuma, Shaokun Zhang and 8 more
arXiv
2609.39982 · PDF
Citations
Not counted yet · Semantic Scholar
Upvotes
112 · Hugging Face
Lab
NVIDIA · on Companies · on Acquisitions · on Quarterly · on Paydays · on TechConf

Changes

What changed
New paperFound by the weekly scan, unverifiedOct 5, 2026today

New papers by email

Monday afternoons, only in weeks with new papers from the labs.

Double opt-in. Unsubscribe any time.