Skip to content

Dynamic Chunking Diffusion Transformer research paper by AMD, 2026

AMD · Mar 6, 2026 · Architectures · 2 citations · 16 upvotes · unverified 6 months ago

Read on arXiv

What it shows

Dynamic Chunking Diffusion Transformer adapts token sequence length based on image content and diffusion timestep, improving efficiency and performance over fixed-token approaches.

UnverifiedHugging Face's summary; not yet checked by hand.

More from AMD

All 6
PaperCitations
Instella-MoE Technical ReportInstella-MoE is an open Mixture-of-Experts language model trained on AMD GPUs using sparse activation, Gated Multi-head Latent Attention, and multi-stage post-training to achieve strong benchmark performance with full reproducibility.Foundation models · Sep 2026 · Unverified3 weeks ago0
Stabilizing Efficient Reasoning with Step-Level Advantage SelectionShort-context post-training induces reasoning compression but causes instability; Step-level Advantage Selection addresses this by selectively adjusting reasoning steps based on confidence and verification outcomes, improving accuracy-efficiency trade-off in reasoning tasks.Reasoning · Apr 2026 · Unverified5 months ago1
DUET-VLM: Dual stage Unified Efficient Token reduction for VLM Training and InferenceDUET-VLM presents a dual compression framework that reduces visual tokens while maintaining high accuracy in vision-language models through coordinated vision and language backbone processing.Inference and efficiency · Feb 2026 · Unverified7 months ago0
Instella: Fully Open Language Models with Stellar PerformanceInstella, a family of fully open large language models, achieves state-of-the-art performance using open data and is competitive with leading open-weight models, with specialized variants for long context and mathematical reasoning.Alignment and safety · Nov 2025 · Unverified10 months ago6
XModBench: Benchmarking Cross-Modal Capabilities and Consistency in Omni-Language ModelsXModBench evaluates cross-modal consistency in OLLMs, revealing challenges in spatial and temporal reasoning, modality disparities, and directional imbalance.Agents and evaluation · Oct 2025 · Unverified11 months ago4
Topic
PaperCitations
Attention Is All You NeedIntroduced the Transformer, the attention-only architecture behind nearly every large language model since.Google · Jun 20179 years ago194k
Mamba: Linear-Time Sequence Modeling with Selective State SpacesA selective state space model that scales linearly with sequence length and matches Transformers on language.Carnegie Mellon University · Dec 20232 years ago9,252
CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationCodeT5, a unified encoder-decoder Transformer model, improves code understanding and generation by leveraging semantic information from identifiers and user comments.Salesforce · Sep 20215 years ago2,510
VGGT: Visual Geometry Grounded TransformerVGGT, a feed-forward neural network, efficiently infers multiple 3D attributes from single or multiple views, outperforming alternatives and enhancing downstream tasks without post-processing.Meta · Mar 2025 · Unverified1 year ago1,807
DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language ModelsThe DeepSeekMoE architecture improves expert specialization in Mixture-of-Experts models by segmenting experts and isolating shared ones, achieving better performance and computational efficiency compared to GShard and other models.DeepSeek · Jan 2024 · Unverified2 years ago1,127
CodeT5+: Open Code Large Language Models for Code Understanding and GenerationCodeT5+, a family of flexible encoder-decoder LLMs initialized with frozen LLMs and trained with a mixture of pretraining objectives, achieves state-of-the-art performance across various code-related tasks.Salesforce · May 20233 years ago792
About this paper
Authors
Akash Haridas, Utkarsh Saxena, Parsa Ashrafi Fashi and 3 more
arXiv
2603.06351 · PDF
Citations
2, 0 influential · Semantic Scholar
Upvotes
16 · Hugging Face
Lab
AMD · on Companies · on Acquisitions · on Quarterly · on Paydays · on TechConf

Changes

What changed
Influential citationsfirst count: 0Sep 28, 2026today
Citationsfirst count: 2Sep 28, 2026today
New paperFound by the weekly scan, unverifiedSep 28, 2026today

New papers by email

Monday afternoons, only in weeks with new papers from the labs.

Double opt-in. Unsubscribe any time.