Skip to content

Grounding Computer Use Agents on Human Demonstrations research paper by ServiceNow, 2025

ServiceNow · Nov 10, 2025 · Agents and evaluation · 107 upvotes 10 months ago

Read on arXiv

What it shows

GroundCUA, a large-scale desktop grounding dataset, enables the development of GroundNext models that achieve state-of-the-art performance in mapping instructions to UI elements with less training data.

Hugging Face's summary; not yet checked by hand.

More from ServiceNow

All 20
PaperCitations
AgentJudgeBench: A Multi-Difficulty Benchmark for Evaluating LLM Judges on Agentic Tool-CallingAgentJudgeBench reveals that LLM judges face structural reliability limits on dependency-driven agentic tool-calling workflows, with alignment degrading by difficulty and ground-truth exposure yielding mixed effects.Agents and evaluation · Aug 2026 · Unverified4 weeks ago-
StarHarness: Evolving Harnesses with Stratified Search for Enterprise EnvironmentsStarHarness evolves fixed-weight agent harnesses via stratified task pools and hidden selection to improve enterprise tool-use performance and cross-model transfer.Retrieval and data · Aug 2026 · Unverified4 weeks ago-
SynthDocBench: Controlled Benchmark for Long-Context Visual Document UnderstandingSynthDocBench is a synthetic long-context visual document benchmark that isolates failure modes in vision-language models, revealing sharp length degradation, positional sensitivity, and chart comprehension breakdowns.Agents and evaluation · Jul 2026 · Unverified2 months ago-
EVA-Bench: A New End-to-end Framework for Evaluating Voice AgentsEVA-Bench presents a comprehensive evaluation framework for voice agents that simulates realistic conversations and measures performance across multiple voice-specific failure modes using novel accuracy and experience metrics.Agents and evaluation · May 2026 · Unverified4 months ago-
Do Enterprise Systems Need Learned World Models? The Importance of Context to Infer DynamicsEnterprise discovery agents that read system configuration at runtime outperform traditional world models in configurable environments where dynamics change over time.Agents and evaluation · May 2026 · Unverified4 months ago-
Apriel-Reasoner: RL Post-Training for General-Purpose and Efficient ReasoningApriel-Reasoner is a 15B-parameter language model trained with reproducible multi-domain reinforcement learning to improve reasoning efficiency and accuracy across diverse tasks while reducing inference costs.Reasoning · Apr 20265 months ago-
Therefore I am. I ThinkReasoning models appear to encode action choices before beginning textual deliberation, as evidenced by early decision detection and activation steering effects.Reasoning · Apr 2026 · Unverified5 months ago-
Terminal Agents Suffice for Enterprise AutomationSimple terminal-based coding agents using programmatic interfaces and foundation models can effectively perform enterprise tasks comparable to or better than complex tool-augmented agents.Agents and evaluation · Mar 2026 · Unverified5 months ago-
Topic
PaperCitations
Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic CapabilitiesGemini 2.X model family, including Gemini 2.5 Pro and Flash, offers superior coding, reasoning, and multimodal understanding capabilities across a range of computational efficiencies.Google · Jul 2025 · Unverified1 year ago4,134
DeepSeek-V3.2: Pushing the Frontier of Open Large Language ModelsDeepSeek-V3.2 introduces DeepSeek Sparse Attention and a scalable reinforcement learning framework, achieving superior reasoning and performance compared to GPT-5 and Gemini-3.0-Pro in complex reasoning tasks.DeepSeek · Dec 2025 · Unverified9 months ago761
WebWatcher: Breaking New Frontier of Vision-Language Deep Research AgentWebWatcher, a multimodal agent with enhanced visual-language reasoning, outperforms existing agents in complex visual and textual information retrieval tasks using synthetic trajectories and reinforcement learning.Alibaba (Qwen) · Aug 2025 · Unverified1 year ago117
SkillOpt: Executive Strategy for Self-Evolving Agent SkillsSkillOpt introduces a systematic text-space optimizer for agent skills that trains skills as external agent state with stable updates and zero deployment inference overhead, achieving superior performance across multiple benchmarks and execution environments.Microsoft · May 2026 · Unverified4 months ago85
Nemotron 3 Nano: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic ReasoningWe present Nemotron 3 Nano 30B-A3B, a Mixture-of-Experts hybrid Mamba-Transformer language model.NVIDIA · Dec 2025 · Unverified9 months ago81
AgentFold: Long-Horizon Web Agents with Proactive Context ManagementAgentFold, a novel proactive context management paradigm, enhances long-horizon task performance through dynamic context folding, achieving superior results on benchmarks compared to larger models and proprietary agents.Alibaba (Qwen) · Oct 2025 · Unverified11 months ago77
About this paper
Authors
Aarash Feizi, Shravan Nayak, Xiangru Jian and 14 more
arXiv
2511.07332 · PDF
Citations
Not counted yet · Semantic Scholar
Upvotes
107 · Hugging Face
Code
github.com/ServiceNow/GroundCUA
Lab
ServiceNow · on Companies · on Acquisitions · on Quarterly · on Paydays · on TechConf

Changes

What changed
New paperAdded to the listSep 26, 2026today

New papers by email

Monday afternoons, only in weeks with new papers from the labs.

Double opt-in. Unsubscribe any time.