Skip to content

Decoding Looped Transformers Better for (Almost) Free research paper by Apple, 2026

Apple · Oct 1, 2026 · Training and scaling · 43 upvotes · unverified 4 days ago

Read on arXiv

What it shows

Looped Transformers achieve parameter efficiency by repeatedly executing a shared block across recurrent loops.

By Weihao Liu, Huangjie Zheng, Tianrong Chen and 5 more · arXiv 2610.02185 · PDF

UnverifiedHugging Face's summary; not yet checked by hand.

More from Apple

All 10
PaperCitations
MintAct: A Unified Visual Agent for Digital EnvironmentsWe present MintAct, a family of vision-language models that unifies UI grounding, multi-step navigation across mobile, desktop, and web, and visual tool use, trained at 2B, 4B, and 8B scales.Agents and evaluation · Sep 2026 · Unverified2 weeks ago0
It Takes Two to Match: Co-Evolving Generative Retriever with Reinforcement LearningCoGR trains LLMs to generate compact keywords for both queries and items, enabling direct inverted-index retrieval optimized via co-evolving reinforcement learning.Retrieval and data · Sep 2026 · Unverified4 weeks ago1
Embarrassingly Simple Self-Distillation Improves Code GenerationSimple self-distillation improves code generation in large language models by fine-tuning on model-generated samples, effectively addressing precision-exploration trade-offs in decoding.Foundation models · Apr 2026 · Unverified6 months ago34
CHIMERA: Compact Synthetic Data for Generalizable LLM ReasoningA synthetic reasoning dataset called CHIMERA is introduced to overcome data-centric challenges in training large language models for cross-domain reasoning, achieving performance comparable to much larger models.Reasoning · Mar 2026 · Unverified7 months ago2
Sharp Monocular View Synthesis in Less Than a SecondSHARP synthesizes photorealistic views from a single image using a 3D Gaussian representation, achieving state-of-the-art results with rapid processing.Multimodal and robotics · Dec 2025 · Unverified9 months ago25
One Layer Is Enough: Adapting Pretrained Visual Encoders for Image GenerationFAE, a framework using a feature auto-encoder and dual decoders, adapts pre-trained visual representations for generative models, achieving high performance in image generation tasks.Multimodal and robotics · Dec 2025 · Unverified10 months ago28
STARFlow-V: End-to-End Video Generative Modeling with Normalizing FlowSTARFlow-V, a normalizing flow-based video generator, offers end-to-end learning, robust causal prediction, and high-quality video generation with practical sampling efficiency.Multimodal and robotics · Nov 2025 · Unverified10 months ago9
CLaRa: Bridging Retrieval and Generation with Continuous Latent ReasoningCLaRa enhances retrieval-augmented generation by introducing unified embedding-based compression and joint optimization, achieving state-of-the-art performance in QA benchmarks.Reasoning · Nov 2025 · Unverified10 months ago12
Topic
PaperCitations
LoRA: Low-Rank Adaptation of Large Language ModelsLoRA fine-tunes a large model by training small low-rank matrices, cutting trainable parameters by 10,000 times.Microsoft · Jun 20215 years ago23.8k
Scaling Laws for Neural Language ModelsLanguage model loss falls as a smooth power law as model size, data and compute grow.OpenAI · Jan 20206 years ago9,190
Training Compute-Optimal Large Language ModelsChinchilla: for a fixed compute budget, train a smaller model on more data; parameters and tokens should grow together.Google DeepMind · Mar 20224 years ago3,835
Megatron-LM: Training Multi-Billion Parameter Language Models Using Model ParallelismMegatron-LM splits each Transformer layer across GPUs to train models with billions of parameters.NVIDIA · Sep 20197 years ago3,236
DeepSeek LLM: Scaling Open-Source Language Models with LongtermismDeepSeek LLM, an open-source language model project, develops a large dataset and employs SFT and DPO to achieve performance surpassing LLaMA-2 70B and GPT-3.5 in various benchmarks and open-ended evaluations.DeepSeek · Jan 2024 · Unverified2 years ago861
LoRA Learns Less and Forgets LessLoRA learns less than full fine-tuning on code and math, but forgets less of what the model already knew.Databricks · May 20242 years ago422
About this paper
Authors
Weihao Liu, Huangjie Zheng, Tianrong Chen and 5 more
arXiv
2610.02185 · PDF
Citations
Not counted yet · Semantic Scholar
Upvotes
43 · Hugging Face
Lab
Apple · on Companies · on Quarterly · on Paydays · on TechConf

Changes

What changed
New paperFound by the weekly scan, unverifiedOct 5, 2026today

New papers by email

Monday afternoons, only in weeks with new papers from the labs.

Double opt-in. Unsubscribe any time.