Skip to content
Papers.

PhyCritic: Multimodal Critic Models for Physical AI research paper by NVIDIA, 2026

NVIDIA · Feb 11, 2026 · Multimodal and robotics · 11 citations · 55 upvotes · unverified

Read on arXiv

What it shows

PhyCritic is a multimodal critic model designed for physical AI tasks through a two-stage RLVR pipeline that enhances perception and reasoning capabilities.

UnverifiedHugging Face's summary; not yet checked by hand.

More from NVIDIA

All 75
PaperCitations
SoL-Pi: Recursively Scaling Auto-Research Loops for Efficient Agent HarnessAs coding agents move from supervised code completion to unattended, around-the-clock exploration, their work expands from isolated predictions into long trajectories of reasoning, tool use, and feedback.Agents and evaluation · Sep 2026 · Unverified1
An Open Recipe for IMO Gold: Training Nemotron for Olympiad MathematicsA natural-language proof-generation pipeline using post-trained Nemotron 3 Ultra checkpoints achieves gold-medal performance on IMO 2026 through iterative verification and refinement without external tools.Training and scaling · Sep 2026 · Unverified0
Online Draft Co-Training for Speculative Decoding in Large-Scale, Long-Context RL Post-TrainingA system for large-scale online draft co-training accelerates speculative decoding in RL post-training by extending context-parallel attention and adding cross-stage feature transport.Training and scaling · Sep 2026 · Unverified1
Sol-Attn: Accelerating Video Generation Inference via On-the-Fly Attention SparsificationSol-Attn improves training-free sparse attention for diffusion transformers by combining dynamic block routing, sparse computation, and approximation correction in a single online pass to accelerate video generation without sacrificing quality.Inference and efficiency · Jul 2026 · Unverified2
SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video GenerationSANA-Video 2.0 is a hybrid video diffusion transformer that combines linear and softmax attention to generate high-resolution video efficiently on a single GPU.Inference and efficiency · Jul 2026 · Unverified2
Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement LearningMolt is a compact PyTorch framework for agentic reinforcement learning that enables efficient asynchronous training of multimodal and mixture-of-experts policies with minimal overhead.Agents and evaluation · Jul 2026 · Unverified1
NVIDIA-labs OO Agents: Native Python Object-Oriented AgentsNOOA treats AI agents as Python objects whose methods and fields define actions, state, and prompts, enabling deterministic testing and LLM-driven runtime completion within a unified programming model.Agents and evaluation · Jul 2026 · Unverified0
ASPIRE: Agentic /Skills Discovery for RoboticsASPIRE is a continual learning system that autonomously develops and refines robot control programs through iterative exploration, achieving superior performance and zero-shot generalization in manipulation and household tasks while enabling sim-to-real transfer.Agents and evaluation · Jun 2026 · Unverified21
Topic
PaperCitations
Segment AnythingA promptable model and a dataset of over a billion masks that cut out any object in any image.Meta · Apr 202315.7k
GPT-4o System CardGPT-4o is an omnimodal autoregressive model trained to handle text, audio, image, and video inputs, offering high-performance outputs across these modalities, with particular strengths in vision and audio.OpenAI · Oct 2024 · Unverified4,980
SAM 3: Segment Anything with ConceptsSegment Anything Model 3 achieves state-of-the-art performance in promptable concept segmentation and tracking by leveraging a unified model architecture with decoupled recognition and localization.Meta · Nov 2025 · Unverified999
DeepSeek-VL: Towards Real-World Vision-Language UnderstandingDeepSeek-VL is an open-source vision-language model that achieves state-of-the-art performance in real-world applications by combining a hybrid vision encoder with effective pretraining strategies to preserve language model capabilities.DeepSeek · Mar 2024 · Unverified889
Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and GenerationJanus, an autoregressive framework with separate visual encoding pathways within a unified transformer architecture, enhances performance in unified multimodal understanding and generation.DeepSeek · Oct 2024 · Unverified481
SAM 3D: 3Dfy Anything in ImagesSAM 3D is a generative model that reconstructs 3D objects from single images using a multi-stage training framework that includes synthetic pretraining and real-world alignment, achieving high performance in human preference tests.Meta · Nov 2025 · Unverified256
About this paper
Authors
Tianyi Xiong, Shihao Wang, Guilin Liu and 5 more
arXiv
2602.11124 · PDF
Venue
arXiv.org
Citations
11, 1 influential · Semantic Scholar
Upvotes
55 · Hugging Face
Lab
NVIDIA · on Companies · on Acquisitions · on Quarterly · on Paydays · on TechConf

Changes

What changed
Influential citationsfirst count: 1Sep 25, 2026
Citationsfirst count: 11Sep 25, 2026
New paperFound by the weekly scan, unverifiedSep 25, 2026

Sources: each lab's own papers and arXiv, with citation and upvote counts from Semantic Scholar and Hugging Face. One-line summaries are for orientation, not a substitute for the paper. Logos via logo.dev; trademarks belong to their owners.

New papers by email

Monday afternoons, only in weeks with new papers from the labs.

Double opt-in. Unsubscribe any time.