Skip to content
Papers.

MARS: Modular Agent with Reflective Search for Automated AI Research research paper by Google, 2026

Google · Feb 2, 2026 · Agents and evaluation · 22 citations · 67 upvotes · unverified

Read on arXiv

What it shows

MARS is a modular AI research automation framework that uses budget-aware planning, modular construction, and reflective memory to achieve state-of-the-art performance in autonomous machine learning research.

UnverifiedHugging Face's summary; not yet checked by hand.

More from Google

All 33
PaperCitations
RRSI: Regularized Recursive Self-Improvement of Agent HarnessesRRSI keeps self-improving agent harnesses from memorising their training tasks, so gains carry over to new benchmarks.Agents and evaluation · Sep 20260
Verifiable Social Reasoning for LLM AssistantsLLM assistants are widely used for daily social advice, yet evaluating their social reasoning in such consultation settings remains challenging since (i) it requires setups where the assistant learns about social...Reasoning · Sep 2026 · Unverified0
Dream-RSI: Recursive Self-Improvement through Evolving WorldsDream-RSI enables scalable recursive self-improvement by using historical discovery replay to evaluate exploration policies offline, reducing costly online evaluations.Retrieval and data · Sep 2026 · Unverified0
Procedural Graphs: Self-Evolving Execution Structures for LLM AgentsA procedural graph framework organizes agent actions into structured relational triplets, providing situational guidance and self-evolving topology to improve long-horizon tool use.Foundation models · Sep 2026 · Unverified0
WikiSkill: Compiling Agent Experience into Persistent Knowledge for Skill EvolutionWikiSkill co-evolves reusable agent skills with a persistent knowledge base to systematically accumulate experience and improve performance across models.Agents and evaluation · Aug 2026 · Unverified0
EnvHarness: Awakening Static Worlds for Agent LearningEnvHarness and EnvRigger dynamically reshape static environments via programmable plugins to target agent weaknesses and improve reinforcement learning co-evolution.Agents and evaluation · Aug 2026 · Unverified2
Concurrent Image Understanding and Generation: Self-Correcting Coupled Markov Jump ProcessesA coupled Markov jump process with cross-modal attention and remasking enables a training-free single-pass sampler for joint multimodal generation that improves with more denoising steps.Multimodal and robotics · Jul 2026 · Unverified1
Reinforcement Learning with Metacognitive Feedback Elicits Faithful Uncertainty Expression in LLMsReinforcement learning with metacognitive feedback and metacognitive data selection improve large language model calibration by enabling accurate self-assessment of performance and uncertainty.Foundation models · Jun 2026 · Unverified1
Topic
PaperCitations
DeepSeek-V3.2: Pushing the Frontier of Open Large Language ModelsDeepSeek-V3.2 introduces DeepSeek Sparse Attention and a scalable reinforcement learning framework, achieving superior reasoning and performance compared to GPT-5 and Gemini-3.0-Pro in complex reasoning tasks.DeepSeek · Dec 2025 · Unverified761
WebWatcher: Breaking New Frontier of Vision-Language Deep Research AgentWebWatcher, a multimodal agent with enhanced visual-language reasoning, outperforms existing agents in complex visual and textual information retrieval tasks using synthetic trajectories and reinforcement learning.Alibaba (Qwen) · Aug 2025 · Unverified117
SkillOpt: Executive Strategy for Self-Evolving Agent SkillsSkillOpt introduces a systematic text-space optimizer for agent skills that trains skills as external agent state with stable updates and zero deployment inference overhead, achieving superior performance across multiple benchmarks and execution environments.Microsoft · May 2026 · Unverified85
Nemotron 3 Nano: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic ReasoningWe present Nemotron 3 Nano 30B-A3B, a Mixture-of-Experts hybrid Mamba-Transformer language model.NVIDIA · Dec 2025 · Unverified81
AgentFold: Long-Horizon Web Agents with Proactive Context ManagementAgentFold, a novel proactive context management paradigm, enhances long-horizon task performance through dynamic context folding, achieving superior results on benchmarks compared to larger models and proprietary agents.Alibaba (Qwen) · Oct 2025 · Unverified77
Agent Learning via Early ExperienceEarly experience, using agent-generated interaction data without reward signals, improves policy effectiveness and generalization, serving as a bridge between imitation learning and reinforcement learning.Meta · Oct 2025 · Unverified64
About this paper
Authors
Jiefeng Chen, Bhavana Dalvi Mishra, Jaehyun Nam and 3 more
arXiv
2602.02660 · PDF
Venue
arXiv.org
Citations
22, 4 influential · Semantic Scholar
Upvotes
67 · Hugging Face
Code
github.com/jfc43/MARS
Lab
Google · on Companies · on Acquisitions · on Paydays · on TechConf · on Releases

Changes

What changed
Influential citationsfirst count: 4Sep 25, 2026
Citationsfirst count: 22Sep 25, 2026
New paperFound by the weekly scan, unverifiedSep 25, 2026

Sources: each lab's own papers and arXiv, with citation and upvote counts from Semantic Scholar and Hugging Face. One-line summaries are for orientation, not a substitute for the paper. Logos via logo.dev; trademarks belong to their owners.

New papers by email

Monday afternoons, only in weeks with new papers from the labs.

Double opt-in. Unsubscribe any time.