Skip to content
Papers.

Don't Mask the Environment: Observation Supervision Changes How Agents Explore Under RL research paper by Amazon, 2026

Amazon · Sep 17, 2026 · Agents and evaluation · 0 citations · 42 upvotes · unverified

Read on arXiv

What it shows

Agent trajectories record what an agent does and what happens next.

UnverifiedHugging Face's summary; not yet checked by hand.

More from Amazon

All 6
PaperCitations
VIDEOP2R: Video Understanding from Perception to ReasoningVideoP2R, a process-aware reinforcement fine-tuning framework, improves video reasoning and understanding by modeling perception and reasoning separately, achieving state-of-the-art results on multiple benchmarks.Reasoning · Nov 2025 · Unverified9
Adaptive Multi-Agent Response Refinement in Conversational SystemsA multi-agent framework enhances conversational quality by refining responses through agents responsible for factuality, personalization, and coherence, outperforming existing methods on challenging datasets.Agents and evaluation · Nov 2025 · Unverified3
Chronos-2: From Univariate to Universal ForecastingChronos-2, a pretrained model with a group attention mechanism, achieves state-of-the-art performance in zero-shot univariate, multivariate, and covariate-informed forecasting tasks.Inference and efficiency · Oct 2025 · Unverified178
When Thoughts Meet Facts: Reusable Reasoning for Long-Context LMsThought templates enhance long-context language models by structuring evidence combination and guiding multi-hop inference, leading to consistent performance improvements across various benchmarks.Reasoning · Oct 2025 · Unverified3
TaTToo: Tool-Grounded Thinking PRM for Test-Time Scaling in Tabular ReasoningTaTToo, a novel table-grounded Process Reward Model, enhances tabular reasoning by explicitly addressing table-specific operations and integrating tool-based verification, leading to significant performance improvements over existing PRMs.Reasoning · Oct 2025 · Unverified12
Topic
PaperCitations
Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic CapabilitiesGemini 2.X model family, including Gemini 2.5 Pro and Flash, offers superior coding, reasoning, and multimodal understanding capabilities across a range of computational efficiencies.Google · Jul 2025 · Unverified4,134
DeepSeek-V3.2: Pushing the Frontier of Open Large Language ModelsDeepSeek-V3.2 introduces DeepSeek Sparse Attention and a scalable reinforcement learning framework, achieving superior reasoning and performance compared to GPT-5 and Gemini-3.0-Pro in complex reasoning tasks.DeepSeek · Dec 2025 · Unverified761
WebWatcher: Breaking New Frontier of Vision-Language Deep Research AgentWebWatcher, a multimodal agent with enhanced visual-language reasoning, outperforms existing agents in complex visual and textual information retrieval tasks using synthetic trajectories and reinforcement learning.Alibaba (Qwen) · Aug 2025 · Unverified117
SkillOpt: Executive Strategy for Self-Evolving Agent SkillsSkillOpt introduces a systematic text-space optimizer for agent skills that trains skills as external agent state with stable updates and zero deployment inference overhead, achieving superior performance across multiple benchmarks and execution environments.Microsoft · May 2026 · Unverified85
Nemotron 3 Nano: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic ReasoningWe present Nemotron 3 Nano 30B-A3B, a Mixture-of-Experts hybrid Mamba-Transformer language model.NVIDIA · Dec 2025 · Unverified81
AgentFold: Long-Horizon Web Agents with Proactive Context ManagementAgentFold, a novel proactive context management paradigm, enhances long-horizon task performance through dynamic context folding, achieving superior results on benchmarks compared to larger models and proprietary agents.Alibaba (Qwen) · Oct 2025 · Unverified77
About this paper
Authors
Juzheng Zhang, Disha Makhija, Manoj Ghuhan Arivazhagan and 2 more
arXiv
2609.20715 · PDF
Citations
0, 0 influential · Semantic Scholar
Upvotes
42 · Hugging Face
Lab
Amazon · on Companies · on Quarterly · on Paydays

Changes

What changed
Influential citationsfirst count: 0Sep 25, 2026
Citationsfirst count: 0Sep 25, 2026
New paperFound by the weekly scan, unverifiedSep 25, 2026

Sources: each lab's own papers and arXiv, with citation and upvote counts from Semantic Scholar and Hugging Face. One-line summaries are for orientation, not a substitute for the paper. Logos via logo.dev; trademarks belong to their owners.

New papers by email

Monday afternoons, only in weeks with new papers from the labs.

Double opt-in. Unsubscribe any time.