Skip to content
Papers.

Emergent temporal abstractions in autoregressive models enable hierarchical reinforcement learning research paper by Google, 2025

Google · Dec 23, 2025 · Inference and efficiency · 5 citations · 62 upvotes · unverified

Read on arXiv

What it shows

Large-scale autoregressive models pretrained on next-token prediction and finetuned with reinforcement learning (RL) have achieved unprecedented success on many problem domains.

UnverifiedHugging Face's summary; not yet checked by hand.

More from Google

All 33
PaperCitations
RRSI: Regularized Recursive Self-Improvement of Agent HarnessesRRSI keeps self-improving agent harnesses from memorising their training tasks, so gains carry over to new benchmarks.Agents and evaluation · Sep 20260
Verifiable Social Reasoning for LLM AssistantsLLM assistants are widely used for daily social advice, yet evaluating their social reasoning in such consultation settings remains challenging since (i) it requires setups where the assistant learns about social...Reasoning · Sep 2026 · Unverified0
Dream-RSI: Recursive Self-Improvement through Evolving WorldsDream-RSI enables scalable recursive self-improvement by using historical discovery replay to evaluate exploration policies offline, reducing costly online evaluations.Retrieval and data · Sep 2026 · Unverified0
Procedural Graphs: Self-Evolving Execution Structures for LLM AgentsA procedural graph framework organizes agent actions into structured relational triplets, providing situational guidance and self-evolving topology to improve long-horizon tool use.Foundation models · Sep 2026 · Unverified0
WikiSkill: Compiling Agent Experience into Persistent Knowledge for Skill EvolutionWikiSkill co-evolves reusable agent skills with a persistent knowledge base to systematically accumulate experience and improve performance across models.Agents and evaluation · Aug 2026 · Unverified0
EnvHarness: Awakening Static Worlds for Agent LearningEnvHarness and EnvRigger dynamically reshape static environments via programmable plugins to target agent weaknesses and improve reinforcement learning co-evolution.Agents and evaluation · Aug 2026 · Unverified2
Concurrent Image Understanding and Generation: Self-Correcting Coupled Markov Jump ProcessesA coupled Markov jump process with cross-modal attention and remasking enables a training-free single-pass sampler for joint multimodal generation that improves with more denoising steps.Multimodal and robotics · Jul 2026 · Unverified1
Reinforcement Learning with Metacognitive Feedback Elicits Faithful Uncertainty Expression in LLMsReinforcement learning with metacognitive feedback and metacognitive data selection improve large language model calibration by enabling accurate self-assessment of performance and uncertainty.Foundation models · Jun 2026 · Unverified1
Topic
PaperCitations
Efficient Memory Management for Large Language Model Serving with PagedAttentionvLLM manages the KV cache like pages of virtual memory, serving models with 2 to 4 times the throughput.UC Berkeley · Sep 20238,481
FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessFlashAttention computes exact attention with far fewer GPU memory reads and writes, making long sequences faster.Stanford University · May 20225,334
Group Sequence Policy OptimizationGroup Sequence Policy Optimization (GSPO) is a reinforcement learning algorithm that improves training efficiency and performance of large language models by using sequence-level importance ratios and operations.Alibaba (Qwen) · Jul 2025 · Unverified688
Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse AttentionNSA, a trainable sparse attention mechanism, enhances long-context modeling efficiency without sacrificing performance, achieving improvements in speed and accuracy over full attention models.DeepSeek · Feb 2025 · Unverified507
SmolVLM: Redefining small and efficient multimodal modelsSmolVLM, a series of compact multimodal models, achieves high performance with minimal GPU memory usage, making efficient deployment on mobile and edge devices possible.Hugging Face · Apr 2025 · Unverified294
Inference-Time Scaling for Generalist Reward ModelingSelf-Principled Critique Tuning enhances pointwise generative reward modeling for large language models, improving scalability and quality compared to existing methods.DeepSeek · Apr 2025 · Unverified249
About this paper
Authors
Seijin Kobayashi, Yanick Schimpf, Maximilian Schlegel and 12 more
arXiv
2512.20605 · PDF
Venue
arXiv.org
Citations
5, 0 influential · Semantic Scholar
Upvotes
62 · Hugging Face
Lab
Google · on Companies · on Acquisitions · on Paydays · on TechConf · on Releases

Changes

What changed
Influential citationsfirst count: 0Sep 25, 2026
Citationsfirst count: 5Sep 25, 2026
New paperFound by the weekly scan, unverifiedSep 25, 2026

Sources: each lab's own papers and arXiv, with citation and upvote counts from Semantic Scholar and Hugging Face. One-line summaries are for orientation, not a substitute for the paper. Logos via logo.dev; trademarks belong to their owners.

New papers by email

Monday afternoons, only in weeks with new papers from the labs.

Double opt-in. Unsubscribe any time.