Skip to content
Papers.

RLP: Reinforcement as a Pretraining Objective research paper by NVIDIA, 2025

NVIDIA · Sep 26, 2025 · Training and scaling · 23 citations · 47 upvotes · unverified

Read on arXiv

What it shows

RLP, an information-driven reinforcement pretraining objective, enhances reasoning models by integrating exploration into pretraining, leading to significant performance improvements across various benchmarks.

UnverifiedHugging Face's summary; not yet checked by hand.

More from NVIDIA

All 75
PaperCitations
SoL-Pi: Recursively Scaling Auto-Research Loops for Efficient Agent HarnessAs coding agents move from supervised code completion to unattended, around-the-clock exploration, their work expands from isolated predictions into long trajectories of reasoning, tool use, and feedback.Agents and evaluation · Sep 2026 · Unverified1
An Open Recipe for IMO Gold: Training Nemotron for Olympiad MathematicsA natural-language proof-generation pipeline using post-trained Nemotron 3 Ultra checkpoints achieves gold-medal performance on IMO 2026 through iterative verification and refinement without external tools.Training and scaling · Sep 2026 · Unverified0
Online Draft Co-Training for Speculative Decoding in Large-Scale, Long-Context RL Post-TrainingA system for large-scale online draft co-training accelerates speculative decoding in RL post-training by extending context-parallel attention and adding cross-stage feature transport.Training and scaling · Sep 2026 · Unverified1
Sol-Attn: Accelerating Video Generation Inference via On-the-Fly Attention SparsificationSol-Attn improves training-free sparse attention for diffusion transformers by combining dynamic block routing, sparse computation, and approximation correction in a single online pass to accelerate video generation without sacrificing quality.Inference and efficiency · Jul 2026 · Unverified2
SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video GenerationSANA-Video 2.0 is a hybrid video diffusion transformer that combines linear and softmax attention to generate high-resolution video efficiently on a single GPU.Inference and efficiency · Jul 2026 · Unverified2
Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement LearningMolt is a compact PyTorch framework for agentic reinforcement learning that enables efficient asynchronous training of multimodal and mixture-of-experts policies with minimal overhead.Agents and evaluation · Jul 2026 · Unverified1
NVIDIA-labs OO Agents: Native Python Object-Oriented AgentsNOOA treats AI agents as Python objects whose methods and fields define actions, state, and prompts, enabling deterministic testing and LLM-driven runtime completion within a unified programming model.Agents and evaluation · Jul 2026 · Unverified0
ASPIRE: Agentic /Skills Discovery for RoboticsASPIRE is a continual learning system that autonomously develops and refines robot control programs through iterative exploration, achieving superior performance and zero-shot generalization in manipulation and household tasks while enabling sim-to-real transfer.Agents and evaluation · Jun 2026 · Unverified21
Topic
PaperCitations
LoRA: Low-Rank Adaptation of Large Language ModelsLoRA fine-tunes a large model by training small low-rank matrices, cutting trainable parameters by 10,000 times.Microsoft · Jun 202123.3k
Scaling Laws for Neural Language ModelsLanguage model loss falls as a smooth power law as model size, data and compute grow.OpenAI · Jan 20209,065
Training Compute-Optimal Large Language ModelsChinchilla: for a fixed compute budget, train a smaller model on more data; parameters and tokens should grow together.Google DeepMind · Mar 20223,756
DeepSeek LLM: Scaling Open-Source Language Models with LongtermismDeepSeek LLM, an open-source language model project, develops a large dataset and employs SFT and DPO to achieve performance surpassing LLaMA-2 70B and GPT-3.5 in various benchmarks and open-ended evaluations.DeepSeek · Jan 2024 · Unverified855
LoRA Learns Less and Forgets LessLoRA learns less than full fine-tuning on code and math, but forgets less of what the model already knew.Databricks · May 2024407
jina-embeddings-v3: Multilingual Embeddings With Task LoRAjina-embeddings-v3, a large-scale text embedding model, achieves state-of-the-art performance in multilingual and long-context retrieval tasks using Low-Rank Adaptation and Matryoshka Representation Learning.Jina AI · Sep 2024 · Unverified200
About this paper
Authors
Ali Hatamizadeh, Syeda Nahida Akter, Shrimai Prabhumoye and 5 more
arXiv
2510.01265 · PDF
Venue
arXiv.org
Citations
23, 2 influential · Semantic Scholar
Upvotes
47 · Hugging Face
Code
github.com/NVlabs/RLP
Lab
NVIDIA · on Companies · on Acquisitions · on Quarterly · on Paydays · on TechConf

Changes

What changed
Influential citationsfirst count: 2Sep 25, 2026
Citationsfirst count: 23Sep 25, 2026
New paperFound by the weekly scan, unverifiedSep 25, 2026

Sources: each lab's own papers and arXiv, with citation and upvote counts from Semantic Scholar and Hugging Face. One-line summaries are for orientation, not a substitute for the paper. Logos via logo.dev; trademarks belong to their owners.

New papers by email

Monday afternoons, only in weeks with new papers from the labs.

Double opt-in. Unsubscribe any time.