Skip to content

DUET-VLM: Dual stage Unified Efficient Token reduction for VLM Training and Inference research paper by AMD, 2026

AMD · Feb 21, 2026 · Inference and efficiency · 0 citations · 5 upvotes · unverified 7 months ago

Read on arXiv

What it shows

DUET-VLM presents a dual compression framework that reduces visual tokens while maintaining high accuracy in vision-language models through coordinated vision and language backbone processing.

UnverifiedHugging Face's summary; not yet checked by hand.

More from AMD

All 6
PaperCitations
Instella-MoE Technical ReportInstella-MoE is an open Mixture-of-Experts language model trained on AMD GPUs using sparse activation, Gated Multi-head Latent Attention, and multi-stage post-training to achieve strong benchmark performance with full reproducibility.Foundation models · Sep 2026 · Unverified3 weeks ago0
Stabilizing Efficient Reasoning with Step-Level Advantage SelectionShort-context post-training induces reasoning compression but causes instability; Step-level Advantage Selection addresses this by selectively adjusting reasoning steps based on confidence and verification outcomes, improving accuracy-efficiency trade-off in reasoning tasks.Reasoning · Apr 2026 · Unverified5 months ago1
Dynamic Chunking Diffusion TransformerDynamic Chunking Diffusion Transformer adapts token sequence length based on image content and diffusion timestep, improving efficiency and performance over fixed-token approaches.Architectures · Mar 2026 · Unverified6 months ago2
Instella: Fully Open Language Models with Stellar PerformanceInstella, a family of fully open large language models, achieves state-of-the-art performance using open data and is competitive with leading open-weight models, with specialized variants for long context and mathematical reasoning.Alignment and safety · Nov 2025 · Unverified10 months ago6
XModBench: Benchmarking Cross-Modal Capabilities and Consistency in Omni-Language ModelsXModBench evaluates cross-modal consistency in OLLMs, revealing challenges in spatial and temporal reasoning, modality disparities, and directional imbalance.Agents and evaluation · Oct 2025 · Unverified11 months ago4
Topic
PaperCitations
Efficient Memory Management for Large Language Model Serving with PagedAttentionvLLM manages the KV cache like pages of virtual memory, serving models with 2 to 4 times the throughput.UC Berkeley · Sep 20233 years ago8,538
FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessFlashAttention computes exact attention with far fewer GPU memory reads and writes, making long sequences faster.Stanford University · May 20224 years ago5,355
Group Sequence Policy OptimizationGroup Sequence Policy Optimization (GSPO) is a reinforcement learning algorithm that improves training efficiency and performance of large language models by using sequence-level importance ratios and operations.Alibaba (Qwen) · Jul 2025 · Unverified1 year ago694
Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse AttentionNSA, a trainable sparse attention mechanism, enhances long-context modeling efficiency without sacrificing performance, achieving improvements in speed and accuracy over full attention models.DeepSeek · Feb 2025 · Unverified1 year ago508
SmolVLM: Redefining small and efficient multimodal modelsSmolVLM, a series of compact multimodal models, achieves high performance with minimal GPU memory usage, making efficient deployment on mobile and edge devices possible.Hugging Face · Apr 2025 · Unverified1 year ago296
Inference-Time Scaling for Generalist Reward ModelingSelf-Principled Critique Tuning enhances pointwise generative reward modeling for large language models, improving scalability and quality compared to existing methods.DeepSeek · Apr 2025 · Unverified1 year ago249
About this paper
Authors
Aditya Kumar Singh, Hitesh Kandala, Pratik Prabhanjan Brahma and 2 more
arXiv
2602.18846 · PDF
Venue
arXiv.org
Citations
0, 0 influential · Semantic Scholar
Upvotes
5 · Hugging Face
Code
github.com/AMD-AGI/DUET-VLM
Lab
AMD · on Companies · on Acquisitions · on Quarterly · on Paydays · on TechConf

Changes

What changed
Influential citationsfirst count: 0Sep 28, 2026today
Citationsfirst count: 0Sep 28, 2026today
New paperFound by the weekly scan, unverifiedSep 28, 2026today

New papers by email

Monday afternoons, only in weeks with new papers from the labs.

Double opt-in. Unsubscribe any time.