Skip to content
Papers.

Scaling Agent Learning via Experience Synthesis research paper by Meta, 2025

Meta · Nov 5, 2025 · Agents and evaluation · 32 citations · 83 upvotes · unverified

Read on arXiv

What it shows

DreamGym is a unified framework that synthesizes diverse experiences for scalable online RL training, improving agent performance and reducing real-world interactions.

UnverifiedHugging Face's summary; not yet checked by hand.

More from Meta

All 34
PaperCitations
WearableQA: A Benchmark for Health Reasoning over Real-World Wearable DataWearableQA is a benchmark of multiple-choice questions derived from real longitudinal wearable data that evaluates large language model reasoning across data and health dimensions.Agents and evaluation · Sep 2026 · Unverified0
Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and RecipesControlled experiments on when text, image understanding and image generation help or compete when trained together.Multimodal and robotics · Aug 20262
HumanCLAW: Can Vision-Language Models Act Through a Body?HumanCLAW decouples high-level vision-language decisions from low-level motor execution to evaluate embodied action intelligence, revealing that current vision-language models lack embodied self-awareness.Multimodal and robotics · Jul 2026 · Unverified2
TUA-Bench: A Benchmark for General-Purpose Terminal-Use AgentsTUA-Bench presents a comprehensive benchmark for evaluating general-purpose terminal-use agents across diverse digital activities and specialized workflows, revealing significant performance gaps among current frontier agents.Agents and evaluation · Jun 2026 · Unverified2
Think in Strokes, Not Pixels: Process-Driven Image Generation via Interleaved ReasoningProcess-driven image generation decomposes synthesis into iterative steps involving textual planning, visual drafting, textual reflection, and visual refinement, with step-wise supervision ensuring consistency and interpretability.Reasoning · Apr 2026 · Unverified5
V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised LearningV-JEPA 2.1 is a self-supervised model that learns dense visual representations for images and videos through a combination of dense predictive loss, deep self-supervision, multi-modal tokenizers, and effective scaling.Multimodal and robotics · Mar 2026 · Unverified83
ReMix: Reinforcement routing for mixtures of LoRAs in LLM finetuningResearchers address imbalance in routing weights of Mixture-of-LoRAs models by proposing Reinforcement Routing (ReMix), which uses non-learnable weights and reinforcement learning techniques to improve model expressiveness and performance.Inference and efficiency · Mar 2026 · Unverified3
Beyond Language Modeling: An Exploration of Multimodal PretrainingControlled multimodal pretraining experiments reveal key insights about unified visual representations, data complementarity, world modeling emergence, and efficient scaling through mixture-of-experts architectures.Multimodal and robotics · Mar 2026 · Unverified32
Topic
PaperCitations
Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic CapabilitiesGemini 2.X model family, including Gemini 2.5 Pro and Flash, offers superior coding, reasoning, and multimodal understanding capabilities across a range of computational efficiencies.Google · Jul 2025 · Unverified4,134
DeepSeek-V3.2: Pushing the Frontier of Open Large Language ModelsDeepSeek-V3.2 introduces DeepSeek Sparse Attention and a scalable reinforcement learning framework, achieving superior reasoning and performance compared to GPT-5 and Gemini-3.0-Pro in complex reasoning tasks.DeepSeek · Dec 2025 · Unverified761
WebWatcher: Breaking New Frontier of Vision-Language Deep Research AgentWebWatcher, a multimodal agent with enhanced visual-language reasoning, outperforms existing agents in complex visual and textual information retrieval tasks using synthetic trajectories and reinforcement learning.Alibaba (Qwen) · Aug 2025 · Unverified117
SkillOpt: Executive Strategy for Self-Evolving Agent SkillsSkillOpt introduces a systematic text-space optimizer for agent skills that trains skills as external agent state with stable updates and zero deployment inference overhead, achieving superior performance across multiple benchmarks and execution environments.Microsoft · May 2026 · Unverified85
Nemotron 3 Nano: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic ReasoningWe present Nemotron 3 Nano 30B-A3B, a Mixture-of-Experts hybrid Mamba-Transformer language model.NVIDIA · Dec 2025 · Unverified81
AgentFold: Long-Horizon Web Agents with Proactive Context ManagementAgentFold, a novel proactive context management paradigm, enhances long-horizon task performance through dynamic context folding, achieving superior results on benchmarks compared to larger models and proprietary agents.Alibaba (Qwen) · Oct 2025 · Unverified77
About this paper
Authors
Zhaorun Chen, Zhuokai Zhao, Kai Zhang and 15 more
arXiv
2511.03773 · PDF
Venue
arXiv.org
Citations
32, 6 influential · Semantic Scholar
Upvotes
83 · Hugging Face
Lab
Meta · on Companies · on Acquisitions · on Quarterly · on Paydays · on TechConf

Changes

What changed
Influential citationsfirst count: 6Sep 25, 2026
Citationsfirst count: 32Sep 25, 2026
New paperFound by the weekly scan, unverifiedSep 25, 2026

Sources: each lab's own papers and arXiv, with citation and upvote counts from Semantic Scholar and Hugging Face. One-line summaries are for orientation, not a substitute for the paper. Logos via logo.dev; trademarks belong to their owners.

New papers by email

Monday afternoons, only in weeks with new papers from the labs.

Double opt-in. Unsubscribe any time.