Skip to content
Papers.

Nemotron-Cascade: Scaling Cascaded Reinforcement Learning for General-Purpose Reasoning Models research paper by NVIDIA, 2025

NVIDIA · Dec 15, 2025 · Reasoning · 39 citations · 40 upvotes · unverified

Read on arXiv

What it shows

Cascaded domain-wise reinforcement learning (Cascade RL) is proposed to enhance general-purpose reasoning models, achieving state-of-the-art performance across benchmarks and outperforming the teacher model in coding competitions.

UnverifiedHugging Face's summary; not yet checked by hand.

More from NVIDIA

All 75
PaperCitations
SoL-Pi: Recursively Scaling Auto-Research Loops for Efficient Agent HarnessAs coding agents move from supervised code completion to unattended, around-the-clock exploration, their work expands from isolated predictions into long trajectories of reasoning, tool use, and feedback.Agents and evaluation · Sep 2026 · Unverified1
An Open Recipe for IMO Gold: Training Nemotron for Olympiad MathematicsA natural-language proof-generation pipeline using post-trained Nemotron 3 Ultra checkpoints achieves gold-medal performance on IMO 2026 through iterative verification and refinement without external tools.Training and scaling · Sep 2026 · Unverified0
Online Draft Co-Training for Speculative Decoding in Large-Scale, Long-Context RL Post-TrainingA system for large-scale online draft co-training accelerates speculative decoding in RL post-training by extending context-parallel attention and adding cross-stage feature transport.Training and scaling · Sep 2026 · Unverified1
Sol-Attn: Accelerating Video Generation Inference via On-the-Fly Attention SparsificationSol-Attn improves training-free sparse attention for diffusion transformers by combining dynamic block routing, sparse computation, and approximation correction in a single online pass to accelerate video generation without sacrificing quality.Inference and efficiency · Jul 2026 · Unverified2
SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video GenerationSANA-Video 2.0 is a hybrid video diffusion transformer that combines linear and softmax attention to generate high-resolution video efficiently on a single GPU.Inference and efficiency · Jul 2026 · Unverified2
Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement LearningMolt is a compact PyTorch framework for agentic reinforcement learning that enables efficient asynchronous training of multimodal and mixture-of-experts policies with minimal overhead.Agents and evaluation · Jul 2026 · Unverified1
NVIDIA-labs OO Agents: Native Python Object-Oriented AgentsNOOA treats AI agents as Python objects whose methods and fields define actions, state, and prompts, enabling deterministic testing and LLM-driven runtime completion within a unified programming model.Agents and evaluation · Jul 2026 · Unverified0
ASPIRE: Agentic /Skills Discovery for RoboticsASPIRE is a continual learning system that autonomously develops and refines robot control programs through iterative exploration, achieving superior performance and zero-shot generalization in manipulation and household tasks while enabling sim-to-real transfer.Agents and evaluation · Jun 2026 · Unverified21
Topic
PaperCitations
Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsAsking a model to write out its intermediate steps (chain of thought) sharply improves its math and logic answers.Google · Jan 202222.1k
DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language ModelsDeepSeekMath 7B improves mathematical reasoning through enhanced data pre-training and Group Relative Policy Optimization, achieving high scores on MATH benchmark without external tools.DeepSeek · Feb 2024 · Unverified8,994
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement LearningReinforcement learning taught a model long step-by-step reasoning on par with OpenAI o1.DeepSeek · Jan 20255,694
DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code IntelligenceDeepSeek-Coder-V2, a Mixture-of-Experts language model, excels in code-specific tasks by enhancing coding and mathematical reasoning capabilities while expanding language support and context length.DeepSeek · Jun 2024 · Unverified502
MolmoAct: Action Reasoning Models that can Reason in SpaceAction Reasoning Models (ARMs) integrate perception, planning, and control to enable adaptable and explainable robotic behavior, achieving superior performance across various tasks and settings.Ai2 · Aug 2025 · Unverified185
WebShaper: Agentically Data Synthesizing via Information-Seeking FormalizationWebShaper, a formalization-driven framework, synthesizes information-seeking datasets using set theory and Knowledge Projections to enhance reasoning structure and achieve top performance in open-sourced benchmarks.Alibaba (Qwen) · Jul 2025 · Unverified113
About this paper
Authors
Boxin Wang, Chankyu Lee, Nayeon Lee and 9 more
arXiv
2512.13607 · PDF
Venue
arXiv.org
Citations
39, 1 influential · Semantic Scholar
Upvotes
40 · Hugging Face
Lab
NVIDIA · on Companies · on Acquisitions · on Quarterly · on Paydays · on TechConf

Changes

What changed
Influential citationsfirst count: 1Sep 25, 2026
Citationsfirst count: 39Sep 25, 2026
New paperFound by the weekly scan, unverifiedSep 25, 2026

Sources: each lab's own papers and arXiv, with citation and upvote counts from Semantic Scholar and Hugging Face. One-line summaries are for orientation, not a substitute for the paper. Logos via logo.dev; trademarks belong to their owners.

New papers by email

Monday afternoons, only in weeks with new papers from the labs.

Double opt-in. Unsubscribe any time.