Skip to content

WorkArena++: Towards Compositional Planning and Reasoning-based Common Knowledge Work Tasks research paper by ServiceNow, 2024

ServiceNow · Jul 7, 2024 · Reasoning · 2 upvotes 2 years ago

Read on arXiv

What it shows

WorkArena++ benchmark evaluates the effectiveness of LLMs and VLMs in workplace tasks, revealing challenges and providing a mechanism for model fine-tuning.

Hugging Face's summary; not yet checked by hand.

More from ServiceNow

All 20
PaperCitations
AgentJudgeBench: A Multi-Difficulty Benchmark for Evaluating LLM Judges on Agentic Tool-CallingAgentJudgeBench reveals that LLM judges face structural reliability limits on dependency-driven agentic tool-calling workflows, with alignment degrading by difficulty and ground-truth exposure yielding mixed effects.Agents and evaluation · Aug 2026 · Unverified4 weeks ago-
StarHarness: Evolving Harnesses with Stratified Search for Enterprise EnvironmentsStarHarness evolves fixed-weight agent harnesses via stratified task pools and hidden selection to improve enterprise tool-use performance and cross-model transfer.Retrieval and data · Aug 2026 · Unverified4 weeks ago-
SynthDocBench: Controlled Benchmark for Long-Context Visual Document UnderstandingSynthDocBench is a synthetic long-context visual document benchmark that isolates failure modes in vision-language models, revealing sharp length degradation, positional sensitivity, and chart comprehension breakdowns.Agents and evaluation · Jul 2026 · Unverified2 months ago-
EVA-Bench: A New End-to-end Framework for Evaluating Voice AgentsEVA-Bench presents a comprehensive evaluation framework for voice agents that simulates realistic conversations and measures performance across multiple voice-specific failure modes using novel accuracy and experience metrics.Agents and evaluation · May 2026 · Unverified4 months ago-
Do Enterprise Systems Need Learned World Models? The Importance of Context to Infer DynamicsEnterprise discovery agents that read system configuration at runtime outperform traditional world models in configurable environments where dynamics change over time.Agents and evaluation · May 2026 · Unverified4 months ago-
Apriel-Reasoner: RL Post-Training for General-Purpose and Efficient ReasoningApriel-Reasoner is a 15B-parameter language model trained with reproducible multi-domain reinforcement learning to improve reasoning efficiency and accuracy across diverse tasks while reducing inference costs.Reasoning · Apr 20265 months ago-
Therefore I am. I ThinkReasoning models appear to encode action choices before beginning textual deliberation, as evidenced by early decision detection and activation steering effects.Reasoning · Apr 2026 · Unverified5 months ago-
Terminal Agents Suffice for Enterprise AutomationSimple terminal-based coding agents using programmatic interfaces and foundation models can effectively perform enterprise tasks comparable to or better than complex tool-augmented agents.Agents and evaluation · Mar 2026 · Unverified5 months ago-
Topic
PaperCitations
Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsAsking a model to write out its intermediate steps (chain of thought) sharply improves its math and logic answers.Google · Jan 20224 years ago22.1k
DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language ModelsDeepSeekMath 7B improves mathematical reasoning through enhanced data pre-training and Group Relative Policy Optimization, achieving high scores on MATH benchmark without external tools.DeepSeek · Feb 2024 · Unverified2 years ago8,994
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement LearningReinforcement learning taught a model long step-by-step reasoning on par with OpenAI o1.DeepSeek · Jan 20251 year ago5,694
DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code IntelligenceDeepSeek-Coder-V2, a Mixture-of-Experts language model, excels in code-specific tasks by enhancing coding and mathematical reasoning capabilities while expanding language support and context length.DeepSeek · Jun 2024 · Unverified2 years ago502
MolmoAct: Action Reasoning Models that can Reason in SpaceAction Reasoning Models (ARMs) integrate perception, planning, and control to enable adaptable and explainable robotic behavior, achieving superior performance across various tasks and settings.Ai2 · Aug 2025 · Unverified1 year ago185
ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent PlanningThinkAct, a dual-system framework, uses reinforced visual latent planning to enable few-shot adaptation, long-horizon planning, and self-correction in embodied AI tasks by bridging high-level reasoning with low-level action execution.NVIDIA · Jul 2025 · Unverified1 year ago172
About this paper
Authors
Léo Boisvert, Megh Thakkar, Maxime Gasse and 6 more
arXiv
2407.05291 · PDF
Citations
Not counted yet · Semantic Scholar
Upvotes
2 · Hugging Face
Lab
ServiceNow · on Companies · on Acquisitions · on Quarterly · on Paydays · on TechConf

Changes

What changed
New paperAdded to the listSep 26, 2026today

New papers by email

Monday afternoons, only in weeks with new papers from the labs.

Double opt-in. Unsubscribe any time.