Skip to content

Stabilizing Efficient Reasoning with Step-Level Advantage Selection research paper by AMD, 2026

AMD · Apr 27, 2026 · Reasoning · 1 citations · 8 upvotes · unverified 5 months ago

Read on arXiv

What it shows

Short-context post-training induces reasoning compression but causes instability; Step-level Advantage Selection addresses this by selectively adjusting reasoning steps based on confidence and verification outcomes, improving accuracy-efficiency trade-off in reasoning tasks.

UnverifiedHugging Face's summary; not yet checked by hand.

More from AMD

All 6
PaperCitations
Instella-MoE Technical ReportInstella-MoE is an open Mixture-of-Experts language model trained on AMD GPUs using sparse activation, Gated Multi-head Latent Attention, and multi-stage post-training to achieve strong benchmark performance with full reproducibility.Foundation models · Sep 2026 · Unverified3 weeks ago0
Dynamic Chunking Diffusion TransformerDynamic Chunking Diffusion Transformer adapts token sequence length based on image content and diffusion timestep, improving efficiency and performance over fixed-token approaches.Architectures · Mar 2026 · Unverified6 months ago2
DUET-VLM: Dual stage Unified Efficient Token reduction for VLM Training and InferenceDUET-VLM presents a dual compression framework that reduces visual tokens while maintaining high accuracy in vision-language models through coordinated vision and language backbone processing.Inference and efficiency · Feb 2026 · Unverified7 months ago0
Instella: Fully Open Language Models with Stellar PerformanceInstella, a family of fully open large language models, achieves state-of-the-art performance using open data and is competitive with leading open-weight models, with specialized variants for long context and mathematical reasoning.Alignment and safety · Nov 2025 · Unverified10 months ago6
XModBench: Benchmarking Cross-Modal Capabilities and Consistency in Omni-Language ModelsXModBench evaluates cross-modal consistency in OLLMs, revealing challenges in spatial and temporal reasoning, modality disparities, and directional imbalance.Agents and evaluation · Oct 2025 · Unverified11 months ago4
Topic
PaperCitations
Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsAsking a model to write out its intermediate steps (chain of thought) sharply improves its math and logic answers.Google · Jan 20224 years ago22.1k
DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language ModelsDeepSeekMath 7B improves mathematical reasoning through enhanced data pre-training and Group Relative Policy Optimization, achieving high scores on MATH benchmark without external tools.DeepSeek · Feb 2024 · Unverified2 years ago9,044
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement LearningReinforcement learning taught a model long step-by-step reasoning on par with OpenAI o1.DeepSeek · Jan 20251 year ago5,718
DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code IntelligenceDeepSeek-Coder-V2, a Mixture-of-Experts language model, excels in code-specific tasks by enhancing coding and mathematical reasoning capabilities while expanding language support and context length.DeepSeek · Jun 2024 · Unverified2 years ago504
MolmoAct: Action Reasoning Models that can Reason in SpaceAction Reasoning Models (ARMs) integrate perception, planning, and control to enable adaptable and explainable robotic behavior, achieving superior performance across various tasks and settings.Ai2 · Aug 2025 · Unverified1 year ago187
ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent PlanningThinkAct, a dual-system framework, uses reinforced visual latent planning to enable few-shot adaptation, long-horizon planning, and self-correction in embodied AI tasks by bridging high-level reasoning with low-level action execution.NVIDIA · Jul 2025 · Unverified1 year ago174
About this paper
Authors
Han Wang, Xiaodong Yu, Jialian Wu and 4 more
arXiv
2604.24003 · PDF
Venue
Annual Meeting of the Association for Computational Linguistics
Citations
1, 0 influential · Semantic Scholar
Upvotes
8 · Hugging Face
Code
github.com/HanNight/SAS
Lab
AMD · on Companies · on Acquisitions · on Quarterly · on Paydays · on TechConf

Changes

What changed
Influential citationsfirst count: 0Sep 28, 2026today
Citationsfirst count: 1Sep 28, 2026today
New paperFound by the weekly scan, unverifiedSep 28, 2026today

New papers by email

Monday afternoons, only in weeks with new papers from the labs.

Double opt-in. Unsubscribe any time.