Skip to content
Papers.

Hymba: A Hybrid-head Architecture for Small Language Models research paper by NVIDIA, 2024

NVIDIA · Nov 20, 2024 · Architectures · 115 citations · 50 upvotes · unverified

Read on arXiv

What it shows

Hymba, a family of small language models with a hybrid-head architecture combining transformer attention and state space models, achieves state-of-the-art performance with improved efficiency and reduced cache size.

UnverifiedHugging Face's summary; not yet checked by hand.

More from NVIDIA

All 75
PaperCitations
SoL-Pi: Recursively Scaling Auto-Research Loops for Efficient Agent HarnessAs coding agents move from supervised code completion to unattended, around-the-clock exploration, their work expands from isolated predictions into long trajectories of reasoning, tool use, and feedback.Agents and evaluation · Sep 2026 · Unverified1
An Open Recipe for IMO Gold: Training Nemotron for Olympiad MathematicsA natural-language proof-generation pipeline using post-trained Nemotron 3 Ultra checkpoints achieves gold-medal performance on IMO 2026 through iterative verification and refinement without external tools.Training and scaling · Sep 2026 · Unverified0
Online Draft Co-Training for Speculative Decoding in Large-Scale, Long-Context RL Post-TrainingA system for large-scale online draft co-training accelerates speculative decoding in RL post-training by extending context-parallel attention and adding cross-stage feature transport.Training and scaling · Sep 2026 · Unverified1
Sol-Attn: Accelerating Video Generation Inference via On-the-Fly Attention SparsificationSol-Attn improves training-free sparse attention for diffusion transformers by combining dynamic block routing, sparse computation, and approximation correction in a single online pass to accelerate video generation without sacrificing quality.Inference and efficiency · Jul 2026 · Unverified2
SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video GenerationSANA-Video 2.0 is a hybrid video diffusion transformer that combines linear and softmax attention to generate high-resolution video efficiently on a single GPU.Inference and efficiency · Jul 2026 · Unverified2
Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement LearningMolt is a compact PyTorch framework for agentic reinforcement learning that enables efficient asynchronous training of multimodal and mixture-of-experts policies with minimal overhead.Agents and evaluation · Jul 2026 · Unverified1
NVIDIA-labs OO Agents: Native Python Object-Oriented AgentsNOOA treats AI agents as Python objects whose methods and fields define actions, state, and prompts, enabling deterministic testing and LLM-driven runtime completion within a unified programming model.Agents and evaluation · Jul 2026 · Unverified0
ASPIRE: Agentic /Skills Discovery for RoboticsASPIRE is a continual learning system that autonomously develops and refines robot control programs through iterative exploration, achieving superior performance and zero-shot generalization in manipulation and household tasks while enabling sim-to-real transfer.Agents and evaluation · Jun 2026 · Unverified21
Topic
PaperCitations
Attention Is All You NeedIntroduced the Transformer, the attention-only architecture behind nearly every large language model since.Google · Jun 2017194k
Mamba: Linear-Time Sequence Modeling with Selective State SpacesA selective state space model that scales linearly with sequence length and matches Transformers on language.Carnegie Mellon University · Dec 20239,203
VGGT: Visual Geometry Grounded TransformerVGGT, a feed-forward neural network, efficiently infers multiple 3D attributes from single or multiple views, outperforming alternatives and enhancing downstream tasks without post-processing.Meta · Mar 2025 · Unverified1,798
DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language ModelsThe DeepSeekMoE architecture improves expert specialization in Mixture-of-Experts models by segmenting experts and isolating shared ones, achieving better performance and computational efficiency compared to GShard and other models.DeepSeek · Jan 2024 · Unverified1,123
Toto: Time Series Optimized Transformer for ObservabilityToto, a Time Series Optimized Transformer for Observability, achieves state-of-the-art performance in observability and general-purpose forecasting using a vast dataset of time series data.Datadog · Jul 2024 · Unverified43
Physics of Language Models: Part 4.1, Architecture Design and the Magic of Canon LayersControlled synthetic pretraining reveals CANON LAYERS, lightweight architectural components that enhance information flow and model capabilities across various sequence architectures.Meta · Dec 2025 · Unverified40
About this paper
Authors
Xin Dong, Yonggan Fu, Shizhe Diao and 10 more
arXiv
2411.13676 · PDF
Venue
International Conference on Learning Representations
Citations
115, 12 influential · Semantic Scholar
Upvotes
50 · Hugging Face
Code
github.com/NVlabs/hymba
Lab
NVIDIA · on Companies · on Acquisitions · on Quarterly · on Paydays · on TechConf

Changes

What changed
Influential citationsfirst count: 12Sep 25, 2026
Citationsfirst count: 115Sep 25, 2026
New paperFound by the weekly scan, unverifiedSep 25, 2026

Sources: each lab's own papers and arXiv, with citation and upvote counts from Semantic Scholar and Hugging Face. One-line summaries are for orientation, not a substitute for the paper. Logos via logo.dev; trademarks belong to their owners.

New papers by email

Monday afternoons, only in weeks with new papers from the labs.

Double opt-in. Unsubscribe any time.