Skip to content

EnterpriseOps-Gym: Environments and Evaluations for Stateful Agentic Planning and Tool Use in Enterprise Settings research paper by ServiceNow, 2026

ServiceNow · Mar 13, 2026 · Agents and evaluation · 150 upvotes · unverified 6 months ago

Read on arXiv

What it shows

A sandbox of 164 database tables and 512 tools with 1,150 enterprise tasks; the best model tested completed only 37% of them.

UnverifiedSummarised by hand from the abstract.

More from ServiceNow

All 20
PaperCitations
AgentJudgeBench: A Multi-Difficulty Benchmark for Evaluating LLM Judges on Agentic Tool-CallingAgentJudgeBench reveals that LLM judges face structural reliability limits on dependency-driven agentic tool-calling workflows, with alignment degrading by difficulty and ground-truth exposure yielding mixed effects.Agents and evaluation · Aug 2026 · Unverified4 weeks ago-
StarHarness: Evolving Harnesses with Stratified Search for Enterprise EnvironmentsStarHarness evolves fixed-weight agent harnesses via stratified task pools and hidden selection to improve enterprise tool-use performance and cross-model transfer.Retrieval and data · Aug 2026 · Unverified4 weeks ago-
SynthDocBench: Controlled Benchmark for Long-Context Visual Document UnderstandingSynthDocBench is a synthetic long-context visual document benchmark that isolates failure modes in vision-language models, revealing sharp length degradation, positional sensitivity, and chart comprehension breakdowns.Agents and evaluation · Jul 2026 · Unverified2 months ago-
EVA-Bench: A New End-to-end Framework for Evaluating Voice AgentsEVA-Bench presents a comprehensive evaluation framework for voice agents that simulates realistic conversations and measures performance across multiple voice-specific failure modes using novel accuracy and experience metrics.Agents and evaluation · May 2026 · Unverified4 months ago-
Do Enterprise Systems Need Learned World Models? The Importance of Context to Infer DynamicsEnterprise discovery agents that read system configuration at runtime outperform traditional world models in configurable environments where dynamics change over time.Agents and evaluation · May 2026 · Unverified4 months ago-
Apriel-Reasoner: RL Post-Training for General-Purpose and Efficient ReasoningApriel-Reasoner is a 15B-parameter language model trained with reproducible multi-domain reinforcement learning to improve reasoning efficiency and accuracy across diverse tasks while reducing inference costs.Reasoning · Apr 20265 months ago-
Therefore I am. I ThinkReasoning models appear to encode action choices before beginning textual deliberation, as evidenced by early decision detection and activation steering effects.Reasoning · Apr 2026 · Unverified5 months ago-
Terminal Agents Suffice for Enterprise AutomationSimple terminal-based coding agents using programmatic interfaces and foundation models can effectively perform enterprise tasks comparable to or better than complex tool-augmented agents.Agents and evaluation · Mar 2026 · Unverified5 months ago-
Topic
PaperCitations
Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic CapabilitiesGemini 2.X model family, including Gemini 2.5 Pro and Flash, offers superior coding, reasoning, and multimodal understanding capabilities across a range of computational efficiencies.Google · Jul 2025 · Unverified1 year ago4,134
DeepSeek-V3.2: Pushing the Frontier of Open Large Language ModelsDeepSeek-V3.2 introduces DeepSeek Sparse Attention and a scalable reinforcement learning framework, achieving superior reasoning and performance compared to GPT-5 and Gemini-3.0-Pro in complex reasoning tasks.DeepSeek · Dec 2025 · Unverified9 months ago761
WebWatcher: Breaking New Frontier of Vision-Language Deep Research AgentWebWatcher, a multimodal agent with enhanced visual-language reasoning, outperforms existing agents in complex visual and textual information retrieval tasks using synthetic trajectories and reinforcement learning.Alibaba (Qwen) · Aug 2025 · Unverified1 year ago117
SkillOpt: Executive Strategy for Self-Evolving Agent SkillsSkillOpt introduces a systematic text-space optimizer for agent skills that trains skills as external agent state with stable updates and zero deployment inference overhead, achieving superior performance across multiple benchmarks and execution environments.Microsoft · May 2026 · Unverified4 months ago85
Nemotron 3 Nano: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic ReasoningWe present Nemotron 3 Nano 30B-A3B, a Mixture-of-Experts hybrid Mamba-Transformer language model.NVIDIA · Dec 2025 · Unverified9 months ago81
AgentFold: Long-Horizon Web Agents with Proactive Context ManagementAgentFold, a novel proactive context management paradigm, enhances long-horizon task performance through dynamic context folding, achieving superior results on benchmarks compared to larger models and proprietary agents.Alibaba (Qwen) · Oct 2025 · Unverified11 months ago77
About this paper
Authors
Shiva Krishna Reddy Malay, Shravan Nayak, Jishnu Sethumadhavan Nair and 6 more
arXiv
2603.13594 · PDF
Citations
Not counted yet · Semantic Scholar
Upvotes
150 · Hugging Face
Code
github.com/ServiceNow/EnterpriseOps-Gym
Lab
ServiceNow · on Companies · on Acquisitions · on Quarterly · on Paydays · on TechConf

Changes

What changed
New paperFound by the weekly scan, unverifiedSep 26, 2026today

New papers by email

Monday afternoons, only in weeks with new papers from the labs.

Double opt-in. Unsubscribe any time.