Skip to content
Papers.

TurboDiffusion: Accelerating Video Diffusion Models by 100-200 Times research paper by UC Berkeley, 2025

UC Berkeley · Dec 18, 2025 · Multimodal and robotics · 32 citations · 95 upvotes · unverified

Read on arXiv

What it shows

TurboDiffusion accelerates video generation by 100-200x using attention acceleration, step distillation, and quantization, while maintaining video quality.

UnverifiedHugging Face's summary; not yet checked by hand.

More from UC Berkeley

All 12
PaperCitations
FreeToken: Efficient Edge-Native MoE Serving with Bandwidth-Adaptive ExecutionFreeToken is an edge-native Mixture-of-Experts serving system that dynamically maps computation and model state onto heterogeneous local hardware to run large open-weight models on personal machines.Inference and efficiency · Aug 2026 · Unverified2
Playful Agentic Robot LearningEmbodied robots learn reusable skills through self-directed play and exploration, then apply these skills to improve performance on downstream tasks without additional training.Agents and evaluation · Jun 2026 · Unverified7
Agents' Last ExamOver 1,000 real, checkable professional tasks written with 250+ industry experts; the hardest tier is far from solved.Agents and evaluation · Jun 202615
Flash-KMeans: Fast and Memory-Efficient Exact K-MeansFlash-kmeans enables efficient online k-means clustering on GPUs through novel kernel-level optimizations that eliminate I/O bottlenecks and reduce atomic write contention.Inference and efficiency · Mar 2026 · Unverified6
dLLM: Simple Diffusion Language ModelingA unified open-source framework is presented that standardizes core components of diffusion language modeling for reproduction, customization, and accessible development of both large and small models.Foundation models · Feb 2026 · Unverified12
SLA2: Sparse-Linear Attention with Learnable Routing and QATSLA2 improves sparse-linear attention in diffusion models by introducing a learnable router, direct attention formulation, and quantization-aware fine-tuning for enhanced efficiency and quality.Inference and efficiency · Feb 2026 · Unverified20
Quant VideoGen: Auto-Regressive Long Video Generation via 2-Bit KV-Cache QuantizationQuant VideoGen addresses KV cache memory limitations in autoregressive video diffusion models through semantic-aware smoothing and progressive residual quantization, achieving significant memory reduction with minimal latency impact.Inference and efficiency · Feb 2026 · Unverified16
Residual Context Diffusion Language ModelsResidual Context Diffusion (RCD) enhances diffusion large language models by recycling discarded token information through contextual residuals, improving accuracy with minimal computational overhead.Training and scaling · Jan 2026 · Unverified8
Topic
PaperCitations
Segment AnythingA promptable model and a dataset of over a billion masks that cut out any object in any image.Meta · Apr 202315.7k
GPT-4o System CardGPT-4o is an omnimodal autoregressive model trained to handle text, audio, image, and video inputs, offering high-performance outputs across these modalities, with particular strengths in vision and audio.OpenAI · Oct 2024 · Unverified4,980
SAM 3: Segment Anything with ConceptsSegment Anything Model 3 achieves state-of-the-art performance in promptable concept segmentation and tracking by leveraging a unified model architecture with decoupled recognition and localization.Meta · Nov 2025 · Unverified999
DeepSeek-VL: Towards Real-World Vision-Language UnderstandingDeepSeek-VL is an open-source vision-language model that achieves state-of-the-art performance in real-world applications by combining a hybrid vision encoder with effective pretraining strategies to preserve language model capabilities.DeepSeek · Mar 2024 · Unverified889
Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and GenerationJanus, an autoregressive framework with separate visual encoding pathways within a unified transformer architecture, enhances performance in unified multimodal understanding and generation.DeepSeek · Oct 2024 · Unverified481
SAM 3D: 3Dfy Anything in ImagesSAM 3D is a generative model that reconstructs 3D objects from single images using a multi-stage training framework that includes synthetic pretraining and real-world alignment, achieving high performance in human preference tests.Meta · Nov 2025 · Unverified256
About this paper
Authors
Jintao Zhang, Kaiwen Zheng, Kai Jiang and 5 more
arXiv
2512.16093 · PDF
Venue
arXiv.org
Citations
32, 4 influential · Semantic Scholar
Upvotes
95 · Hugging Face
Code
github.com/thu-ml/TurboDiffusion
Lab
UC Berkeley

Changes

What changed
Influential citationsfirst count: 4Sep 25, 2026
Citationsfirst count: 32Sep 25, 2026
New paperFound by the weekly scan, unverifiedSep 25, 2026

Sources: each lab's own papers and arXiv, with citation and upvote counts from Semantic Scholar and Hugging Face. One-line summaries are for orientation, not a substitute for the paper. Logos via logo.dev; trademarks belong to their owners.

New papers by email

Monday afternoons, only in weeks with new papers from the labs.

Double opt-in. Unsubscribe any time.