Skip to content
Papers.

Pico-Banana-400K: A Large-Scale Dataset for Text-Guided Image Editing research paper by Apple, 2025

Apple · Oct 22, 2025 · Retrieval and data · 62 citations · 30 upvotes · unverified

Read on arXiv

What it shows

Pico-Banana-400K is a large-scale, high-quality dataset for instruction-based image editing, featuring diverse edit pairs, multi-turn editing, preference subsets, and long-short instruction pairs, enabling comprehensive research and benchmarking.

UnverifiedHugging Face's summary; not yet checked by hand.

More from Apple

All 9
PaperCitations
MintAct: A Unified Visual Agent for Digital EnvironmentsWe present MintAct, a family of vision-language models that unifies UI grounding, multi-step navigation across mobile, desktop, and web, and visual tool use, trained at 2B, 4B, and 8B scales.Agents and evaluation · Sep 2026 · Unverified0
It Takes Two to Match: Co-Evolving Generative Retriever with Reinforcement LearningCoGR trains LLMs to generate compact keywords for both queries and items, enabling direct inverted-index retrieval optimized via co-evolving reinforcement learning.Retrieval and data · Sep 2026 · Unverified1
Embarrassingly Simple Self-Distillation Improves Code GenerationSimple self-distillation improves code generation in large language models by fine-tuning on model-generated samples, effectively addressing precision-exploration trade-offs in decoding.Foundation models · Apr 2026 · Unverified33
CHIMERA: Compact Synthetic Data for Generalizable LLM ReasoningA synthetic reasoning dataset called CHIMERA is introduced to overcome data-centric challenges in training large language models for cross-domain reasoning, achieving performance comparable to much larger models.Reasoning · Mar 2026 · Unverified2
Sharp Monocular View Synthesis in Less Than a SecondSHARP synthesizes photorealistic views from a single image using a 3D Gaussian representation, achieving state-of-the-art results with rapid processing.Multimodal and robotics · Dec 2025 · Unverified21
One Layer Is Enough: Adapting Pretrained Visual Encoders for Image GenerationFAE, a framework using a feature auto-encoder and dual decoders, adapts pre-trained visual representations for generative models, achieving high performance in image generation tasks.Multimodal and robotics · Dec 2025 · Unverified24
STARFlow-V: End-to-End Video Generative Modeling with Normalizing FlowSTARFlow-V, a normalizing flow-based video generator, offers end-to-end learning, robust causal prediction, and high-quality video generation with practical sampling efficiency.Multimodal and robotics · Nov 2025 · Unverified9
CLaRa: Bridging Retrieval and Generation with Continuous Latent ReasoningCLaRa enhances retrieval-augmented generation by introducing unified embedding-based compression and joint optimization, achieving state-of-the-art performance in QA benchmarks.Reasoning · Nov 2025 · Unverified11
Topic
PaperCitations
From Local to Global: A Graph RAG Approach to Query-Focused SummarizationGraphRAG builds a knowledge graph of a document set so a model can answer questions about the whole collection.Microsoft · Apr 20242,198
DINOv3DINOv3, a self-supervised learning model, achieves superior performance across various vision tasks by scaling datasets and models, addressing dense feature degradation, and enhancing flexibility with post-hoc strategies.Meta · Aug 2025 · Unverified1,498
DeepSeek-Prover: Advancing Theorem Proving in LLMs through Large-Scale Synthetic DataGenerating extensive Lean 4 proof data from mathematical competition problems improved DeepSeekMath 7B's theorem-proving capabilities over GPT-4 and other methods.DeepSeek · May 2024 · Unverified260
Qwen3-VL-Embedding and Qwen3-VL-Reranker: A Unified Framework for State-of-the-Art Multimodal Retrieval and RankingThe Qwen3-VL-Embedding and Qwen3-VL-Reranker models form an end-to-end multimodal search pipeline, leveraging multi-stage training and cross-attention mechanisms to achieve high-precision retrieval across diverse modalities.Alibaba (Qwen) · Jan 2026 · Unverified231
DeepSeek-Prover-V1.5: Harnessing Proof Assistant Feedback for Reinforcement Learning and Monte-Carlo Tree SearchDeepSeek-Prover-V1.5 improves theorem proving by optimizing training and inference, utilizing reinforcement learning, and proposing RMaxTS for diverse proof paths, achieving state-of-the-art results on miniF2F and ProofNet benchmarks.DeepSeek · Aug 2024 · Unverified207
Learning to Discover at Test TimeTest-time training enables AI systems to discover optimal solutions for specific scientific problems through continual learning focused on individual challenges rather than generalization.Stanford University · Jan 2026 · Unverified82
About this paper
Authors
Yusu Qian, Eli Bocek-Rivele, Liangchen Song and 5 more
arXiv
2510.19808 · PDF
Venue
arXiv.org
Citations
62, 5 influential · Semantic Scholar
Upvotes
30 · Hugging Face
Code
github.com/apple/pico-banana-400k
Lab
Apple · on Companies · on Quarterly · on Paydays · on TechConf

Changes

What changed
Influential citationsfirst count: 5Sep 25, 2026
Citationsfirst count: 62Sep 25, 2026
New paperFound by the weekly scan, unverifiedSep 25, 2026

Sources: each lab's own papers and arXiv, with citation and upvote counts from Semantic Scholar and Hugging Face. One-line summaries are for orientation, not a substitute for the paper. Logos via logo.dev; trademarks belong to their owners.

New papers by email

Monday afternoons, only in weeks with new papers from the labs.

Double opt-in. Unsubscribe any time.