Skip to content

AIM: Agentic Idea Management for Automated Research research paper by Google, 2026

Google · Sep 29, 2026 · Agents and evaluation · 50 upvotes · unverified 6 days ago

Read on arXiv

What it shows

Frontier LLMs are increasingly used to automate scientific research through iterative search.

By Hyeong Kyu Choi, Bhavana Dalvi Mishra, Jiefeng Chen and 7 more · arXiv 2609.38445 · PDF

UnverifiedHugging Face's summary; not yet checked by hand.

More from Google

All 36
PaperCitations
RPTune: Learned Context Curation for LLM Catalog SearchFor small merchant businesses (SMBs) whose catalogs fit within a long-context LLM, full-catalog prompting offers a compelling alternative to multi-stage retrieval designed primarily for large marketplaces with...Retrieval and data · Oct 2026 · Unverified4 days ago-
Selecting Diverse SFT Traces Improves Post-RL GeneralizationVerified solutions are not equally useful for preparing reasoning models for reinforcement learning (RL).Reasoning · Sep 2026 · Unverified8 days ago-
RRSI: Regularized Recursive Self-Improvement of Agent HarnessesRRSI keeps self-improving agent harnesses from memorising their training tasks, so gains carry over to new benchmarks.Agents and evaluation · Sep 20262 weeks ago3
Verifiable Social Reasoning for LLM AssistantsLLM assistants are widely used for daily social advice, yet evaluating their social reasoning in such consultation settings remains challenging since (i) it requires setups where the assistant learns about social...Reasoning · Sep 2026 · Unverified2 weeks ago1
Dream-RSI: Recursive Self-Improvement through Evolving WorldsDream-RSI enables scalable recursive self-improvement by using historical discovery replay to evaluate exploration policies offline, reducing costly online evaluations.Retrieval and data · Sep 2026 · Unverified3 weeks ago9
Procedural Graphs: Self-Evolving Execution Structures for LLM AgentsA procedural graph framework organizes agent actions into structured relational triplets, providing situational guidance and self-evolving topology to improve long-horizon tool use.Foundation models · Sep 2026 · Unverified3 weeks ago0
WikiSkill: Compiling Agent Experience into Persistent Knowledge for Skill EvolutionWikiSkill co-evolves reusable agent skills with a persistent knowledge base to systematically accumulate experience and improve performance across models.Agents and evaluation · Aug 2026 · Unverified5 weeks ago6
EnvHarness: Awakening Static Worlds for Agent LearningEnvHarness and EnvRigger dynamically reshape static environments via programmable plugins to target agent weaknesses and improve reinforcement learning co-evolution.Agents and evaluation · Aug 2026 · Unverified6 weeks ago8
Topic
PaperCitations
DeepSeek-V3.2: Pushing the Frontier of Open Large Language ModelsDeepSeek-V3.2 introduces DeepSeek Sparse Attention and a scalable reinforcement learning framework, achieving superior reasoning and performance compared to GPT-5 and Gemini-3.0-Pro in complex reasoning tasks.DeepSeek · Dec 2025 · Unverified10 months ago784
Aya Model: An Instruction Finetuned Open-Access Multilingual Language ModelAya, a multilingual generative language model supporting over 50% lower-resourced languages, excels in both generative and discriminative tasks across 99 languages, and provides extensive evaluations and open-source resources.Cohere · Feb 20242 years ago410
WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?WorkArena and BrowserGym evaluate large language model-based agents' ability to perform enterprise software tasks, revealing gaps in current agent capabilities and differences between open and closed-source LLMs.ServiceNow · Mar 20242 years ago387
Rubrics as Rewards: Reinforcement Learning Beyond Verifiable DomainsRubrics as Rewards (RaR) framework uses structured rubrics as interpretable reward signals for on-policy training, improving performance in real-world reinforcement learning tasks with subjective criteria.Scale AI · Jul 20251 year ago343
SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?SWE-Bench Pro is a challenging benchmark for coding models, featuring complex, enterprise-level problems that require substantial code modifications, with performance evaluations showing significant limitations in current models.Scale AI · Sep 20251 year ago270
A Careful Examination of Large Language Model Performance on Grade School ArithmeticEvaluation of large language models on a new math benchmark, GSM1k, reveals that many models exhibit signs of overfitting to the existing GSM8k benchmark, leading to performance drops on GSM1k.Scale AI · May 20242 years ago242
About this paper
Authors
Hyeong Kyu Choi, Bhavana Dalvi Mishra, Jiefeng Chen and 7 more
arXiv
2609.38445 · PDF
Citations
Not counted yet · Semantic Scholar
Upvotes
50 · Hugging Face
Lab
Google · on Companies · on Acquisitions · on Paydays · on TechConf · on Releases

Changes

What changed
New paperFound by the weekly scan, unverifiedOct 5, 2026today

New papers by email

Monday afternoons, only in weeks with new papers from the labs.

Double opt-in. Unsubscribe any time.