Skip to content
Papers.

SmolVLM: Redefining small and efficient multimodal models research paper by Hugging Face, 2025

Hugging Face · Apr 7, 2025 · Inference and efficiency · 294 citations · 212 upvotes · unverified

Read on arXiv

What it shows

SmolVLM, a series of compact multimodal models, achieves high performance with minimal GPU memory usage, making efficient deployment on mobile and edge devices possible.

UnverifiedHugging Face's summary; not yet checked by hand.

More from Hugging Face

All 3
Topic
PaperCitations
Efficient Memory Management for Large Language Model Serving with PagedAttentionvLLM manages the KV cache like pages of virtual memory, serving models with 2 to 4 times the throughput.UC Berkeley · Sep 20238,481
FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessFlashAttention computes exact attention with far fewer GPU memory reads and writes, making long sequences faster.Stanford University · May 20225,334
Group Sequence Policy OptimizationGroup Sequence Policy Optimization (GSPO) is a reinforcement learning algorithm that improves training efficiency and performance of large language models by using sequence-level importance ratios and operations.Alibaba (Qwen) · Jul 2025 · Unverified688
Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse AttentionNSA, a trainable sparse attention mechanism, enhances long-context modeling efficiency without sacrificing performance, achieving improvements in speed and accuracy over full attention models.DeepSeek · Feb 2025 · Unverified507
Inference-Time Scaling for Generalist Reward ModelingSelf-Principled Critique Tuning enhances pointwise generative reward modeling for large language models, improving scalability and quality compared to existing methods.DeepSeek · Apr 2025 · Unverified249
DeepSeek-OCR: Contexts Optical CompressionDeepSeek-OCR uses optical 2D mapping to compress long contexts, achieving high OCR precision with reduced vision tokens and demonstrating practical value in document processing.DeepSeek · Oct 2025 · Unverified189
About this paper
Authors
Andrés Marafioti, Orr Zohar, Miquel Farré and 14 more
arXiv
2504.05299 · PDF
Venue
arXiv.org
Citations
294, 37 influential · Semantic Scholar
Upvotes
212 · Hugging Face
Lab
Hugging Face · on Companies · on Acquisitions · on Releases

Changes

What changed
Influential citationsfirst count: 37Sep 25, 2026
Citationsfirst count: 294Sep 25, 2026
New paperFound by the weekly scan, unverifiedSep 25, 2026

Sources: each lab's own papers and arXiv, with citation and upvote counts from Semantic Scholar and Hugging Face. One-line summaries are for orientation, not a substitute for the paper. Logos via logo.dev; trademarks belong to their owners.

New papers by email

Monday afternoons, only in weeks with new papers from the labs.

Double opt-in. Unsubscribe any time.