Skip to content
Papers.

Embarrassingly Simple Self-Distillation Improves Code Generation research paper by Apple, 2026

Apple · Apr 1, 2026 · Foundation models · 33 citations · 55 upvotes · unverified

Read on arXiv

What it shows

Simple self-distillation improves code generation in large language models by fine-tuning on model-generated samples, effectively addressing precision-exploration trade-offs in decoding.

UnverifiedHugging Face's summary; not yet checked by hand.

More from Apple

All 9
PaperCitations
MintAct: A Unified Visual Agent for Digital EnvironmentsWe present MintAct, a family of vision-language models that unifies UI grounding, multi-step navigation across mobile, desktop, and web, and visual tool use, trained at 2B, 4B, and 8B scales.Agents and evaluation · Sep 2026 · Unverified0
It Takes Two to Match: Co-Evolving Generative Retriever with Reinforcement LearningCoGR trains LLMs to generate compact keywords for both queries and items, enabling direct inverted-index retrieval optimized via co-evolving reinforcement learning.Retrieval and data · Sep 2026 · Unverified1
CHIMERA: Compact Synthetic Data for Generalizable LLM ReasoningA synthetic reasoning dataset called CHIMERA is introduced to overcome data-centric challenges in training large language models for cross-domain reasoning, achieving performance comparable to much larger models.Reasoning · Mar 2026 · Unverified2
Sharp Monocular View Synthesis in Less Than a SecondSHARP synthesizes photorealistic views from a single image using a 3D Gaussian representation, achieving state-of-the-art results with rapid processing.Multimodal and robotics · Dec 2025 · Unverified21
One Layer Is Enough: Adapting Pretrained Visual Encoders for Image GenerationFAE, a framework using a feature auto-encoder and dual decoders, adapts pre-trained visual representations for generative models, achieving high performance in image generation tasks.Multimodal and robotics · Dec 2025 · Unverified24
STARFlow-V: End-to-End Video Generative Modeling with Normalizing FlowSTARFlow-V, a normalizing flow-based video generator, offers end-to-end learning, robust causal prediction, and high-quality video generation with practical sampling efficiency.Multimodal and robotics · Nov 2025 · Unverified9
CLaRa: Bridging Retrieval and Generation with Continuous Latent ReasoningCLaRa enhances retrieval-augmented generation by introducing unified embedding-based compression and joint optimization, achieving state-of-the-art performance in QA benchmarks.Reasoning · Nov 2025 · Unverified11
Pico-Banana-400K: A Large-Scale Dataset for Text-Guided Image EditingPico-Banana-400K is a large-scale, high-quality dataset for instruction-based image editing, featuring diverse edit pairs, multi-turn editing, preference subsets, and long-short instruction pairs, enabling comprehensive research and benchmarking.Retrieval and data · Oct 2025 · Unverified62
Topic
About this paper
Authors
Ruixiang Zhang, Richard He Bai, Huangjie Zheng and 3 more
arXiv
2604.01193 · PDF
Venue
arXiv.org
Citations
33, 2 influential · Semantic Scholar
Upvotes
55 · Hugging Face
Code
github.com/apple/ml-ssd
Lab
Apple · on Companies · on Quarterly · on Paydays · on TechConf

Changes

What changed
Influential citationsfirst count: 2Sep 25, 2026
Citationsfirst count: 33Sep 25, 2026
New paperFound by the weekly scan, unverifiedSep 25, 2026

Sources: each lab's own papers and arXiv, with citation and upvote counts from Semantic Scholar and Hugging Face. One-line summaries are for orientation, not a substitute for the paper. Logos via logo.dev; trademarks belong to their owners.

New papers by email

Monday afternoons, only in weeks with new papers from the labs.

Double opt-in. Unsubscribe any time.