Skip to content
Papers.

TOPReward: Token Probabilities as Hidden Zero-Shot Rewards for Robotics research paper by Ai2, 2026

Ai2 · Feb 22, 2026 · Inference and efficiency · 28 citations · 26 upvotes · unverified

Read on arXiv

What it shows

TOPReward is a probabilistically grounded temporal value function that uses pretrained video Vision-Language Models to estimate robotic task progress through internal token logits, achieving superior performance in zero-shot evaluations across diverse real-world tasks.

UnverifiedHugging Face's summary; not yet checked by hand.

More from Ai2

All 7
PaperCitations
MolmoMotion: Forecasting Point Trajectories in 3D with Language Instruction3D point motion forecasting model predicts object trajectories from visual history and language goals, demonstrating superior performance on benchmarks and transferring effectively to robot manipulation and video generation tasks.Multimodal and robotics · Jun 2026 · Unverified4
MolmoAct2: Action Reasoning Models for Real-world DeploymentA fully open robot action model, released with new datasets including 720 hours of two-arm teleoperation.Multimodal and robotics · May 202644
WildDet3D: Scaling Promptable 3D Detection in the WildA unified 3D object detection framework with a large-scale dataset enables open-world detection with multiple prompt types and geometric cue integration.Multimodal and robotics · Apr 2026 · Unverified11
Recurrent-Depth VLA: Implicit Test-Time Compute Scaling of Vision-Language-Action Models via Latent Iterative ReasoningRD-VLA introduces a recurrent architecture for vision-language-action models that adapts computational depth through latent iterative refinement, achieving constant memory usage and improved task success rates.Reasoning · Feb 2026 · Unverified16
Olmo 3Olmo 3, a family of state-of-the-art fully-open language models at 7B and 32B parameter scales, excels in long-context reasoning, function calling, coding, instruction following, general chat, and knowledge recall.Reasoning · Dec 2025 · Unverified0
MolmoAct: Action Reasoning Models that can Reason in SpaceAction Reasoning Models (ARMs) integrate perception, planning, and control to enable adaptable and explainable robotic behavior, achieving superior performance across various tasks and settings.Reasoning · Aug 2025 · Unverified185
Topic
PaperCitations
Efficient Memory Management for Large Language Model Serving with PagedAttentionvLLM manages the KV cache like pages of virtual memory, serving models with 2 to 4 times the throughput.UC Berkeley · Sep 20238,481
FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessFlashAttention computes exact attention with far fewer GPU memory reads and writes, making long sequences faster.Stanford University · May 20225,334
Group Sequence Policy OptimizationGroup Sequence Policy Optimization (GSPO) is a reinforcement learning algorithm that improves training efficiency and performance of large language models by using sequence-level importance ratios and operations.Alibaba (Qwen) · Jul 2025 · Unverified688
Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse AttentionNSA, a trainable sparse attention mechanism, enhances long-context modeling efficiency without sacrificing performance, achieving improvements in speed and accuracy over full attention models.DeepSeek · Feb 2025 · Unverified507
SmolVLM: Redefining small and efficient multimodal modelsSmolVLM, a series of compact multimodal models, achieves high performance with minimal GPU memory usage, making efficient deployment on mobile and edge devices possible.Hugging Face · Apr 2025 · Unverified294
Inference-Time Scaling for Generalist Reward ModelingSelf-Principled Critique Tuning enhances pointwise generative reward modeling for large language models, improving scalability and quality compared to existing methods.DeepSeek · Apr 2025 · Unverified249
About this paper
Authors
Shirui Chen, Cole Harrison, Ying-Chun Lee and 6 more
arXiv
2602.19313 · PDF
Venue
arXiv.org
Citations
28, 6 influential · Semantic Scholar
Upvotes
26 · Hugging Face
Code
github.com/TOPReward/TOPReward
Lab
Ai2 · on Companies

Changes

What changed
Influential citationsfirst count: 6Sep 25, 2026
Citationsfirst count: 28Sep 25, 2026
New paperFound by the weekly scan, unverifiedSep 25, 2026

Sources: each lab's own papers and arXiv, with citation and upvote counts from Semantic Scholar and Hugging Face. One-line summaries are for orientation, not a substitute for the paper. Logos via logo.dev; trademarks belong to their owners.

New papers by email

Monday afternoons, only in weeks with new papers from the labs.

Double opt-in. Unsubscribe any time.