Skip to content

ServiceNow research papers

Company · 20 papers · 0 citations · latest Aug 2026 4 weeks ago

Research page

20 of 20 papers, newest first

PaperCitations
AgentJudgeBench: A Multi-Difficulty Benchmark for Evaluating LLM Judges on Agentic Tool-CallingAgentJudgeBench reveals that LLM judges face structural reliability limits on dependency-driven agentic tool-calling workflows, with alignment degrading by difficulty and ground-truth exposure yielding mixed effects.Agents and evaluation · Aug 2026 · Unverified4 weeks ago-
StarHarness: Evolving Harnesses with Stratified Search for Enterprise EnvironmentsStarHarness evolves fixed-weight agent harnesses via stratified task pools and hidden selection to improve enterprise tool-use performance and cross-model transfer.Retrieval and data · Aug 2026 · Unverified4 weeks ago-
SynthDocBench: Controlled Benchmark for Long-Context Visual Document UnderstandingSynthDocBench is a synthetic long-context visual document benchmark that isolates failure modes in vision-language models, revealing sharp length degradation, positional sensitivity, and chart comprehension breakdowns.Agents and evaluation · Jul 2026 · Unverified2 months ago-
EVA-Bench: A New End-to-end Framework for Evaluating Voice AgentsEVA-Bench presents a comprehensive evaluation framework for voice agents that simulates realistic conversations and measures performance across multiple voice-specific failure modes using novel accuracy and experience metrics.Agents and evaluation · May 2026 · Unverified4 months ago-
Do Enterprise Systems Need Learned World Models? The Importance of Context to Infer DynamicsEnterprise discovery agents that read system configuration at runtime outperform traditional world models in configurable environments where dynamics change over time.Agents and evaluation · May 2026 · Unverified4 months ago-
Apriel-Reasoner: RL Post-Training for General-Purpose and Efficient ReasoningApriel-Reasoner is a 15B-parameter language model trained with reproducible multi-domain reinforcement learning to improve reasoning efficiency and accuracy across diverse tasks while reducing inference costs.Reasoning · Apr 20265 months ago-
Therefore I am. I ThinkReasoning models appear to encode action choices before beginning textual deliberation, as evidenced by early decision detection and activation steering effects.Reasoning · Apr 2026 · Unverified5 months ago-
Terminal Agents Suffice for Enterprise AutomationSimple terminal-based coding agents using programmatic interfaces and foundation models can effectively perform enterprise tasks comparable to or better than complex tool-augmented agents.Agents and evaluation · Mar 2026 · Unverified5 months ago-
CUA-Suite: Massive Human-annotated Video Demonstrations for Computer-Use AgentsCUA-Suite introduces a large-scale ecosystem of expert video demonstrations and annotations for computer-use agents, providing continuous screen recordings and detailed reasoning annotations to advance desktop automation capabilities.Agents and evaluation · Mar 20266 months ago-
EnterpriseOps-Gym: Environments and Evaluations for Stateful Agentic Planning and Tool Use in Enterprise SettingsA sandbox of 164 database tables and 512 tools with 1,150 enterprise tasks; the best model tested completed only 37% of them.Agents and evaluation · Mar 2026 · Unverified6 months ago-
Privileged Information Distillation for Language ModelsTraining methods that utilize privileged information for language model distillation in multi-turn environments outperform standard supervised fine-tuning followed by reinforcement learning approaches.Agents and evaluation · Feb 2026 · Unverified7 months ago-
Grounding Computer Use Agents on Human DemonstrationsGroundCUA, a large-scale desktop grounding dataset, enables the development of GroundNext models that achieve state-of-the-art performance in mapping instructions to UI elements with less training data.Agents and evaluation · Nov 202510 months ago-
Apriel-1.5-15b-ThinkerA 15-billion parameter multimodal reasoning model achieves competitive performance through a progressive training methodology without reinforcement learning, demonstrating efficient use of computational resources.Reasoning · Oct 2025 · Unverified12 months ago-
DeepCodeSeek: Real-Time API Retrieval for Context-Aware Code GenerationA novel technique for predicting APIs and generating code in real-time using a compact reranker outperforms larger models with reduced latency, addressing API leaks and unclear usage intent in enterprise code.Retrieval and data · Sep 2025 · Unverified12 months ago-
Optimizing What Matters: AUC-Driven Learning for Robust Neural RetrievalA new training objective, MW loss, is introduced to improve retriever calibration and ranking quality by directly optimizing the Area under the ROC Curve (AUC), outperforming Contrastive Loss in retrieval-augmented generation tasks.Retrieval and data · Sep 2025 · Unverified12 months ago-
How to Train Your LLM Web Agent: A Statistical DiagnosisA study on compute allocation for post-training LLM-based web agents finds that combining supervised fine-tuning with on-policy reinforcement learning improves performance and reduces computational costs compared to using either method alone.Agents and evaluation · Jul 2025 · Unverified1 year ago-
Multi-task retriever fine-tuning for domain-specific and efficient RAGInstruction-tuned retrieval encoder addresses domain-specific challenges for efficient, scalable, and fast Retrieval-Augmented Generation (RAG) applications.Inference and efficiency · Jan 2025 · Unverified1 year ago-
The BrowserGym Ecosystem for Web Agent ResearchBrowserGym provides a standardized environment for evaluating web agents using LLMs, facilitating consistent benchmarking and improving agent development and analysis.Agents and evaluation · Dec 20241 year ago-
WorkArena++: Towards Compositional Planning and Reasoning-based Common Knowledge Work TasksWorkArena++ benchmark evaluates the effectiveness of LLMs and VLMs in workplace tasks, revealing challenges and providing a mechanism for model fine-tuning.Reasoning · Jul 20242 years ago-
WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?WorkArena and BrowserGym evaluate large language model-based agents' ability to perform enterprise software tasks, revealing gaps in current agent capabilities and differences between open and closed-source LLMs.Agents and evaluation · Mar 20242 years ago-
About and links

New papers by email

Monday afternoons, only in weeks with new papers from the labs.

Double opt-in. Unsubscribe any time.