Skip to content
Papers.

Reasoning Models Struggle to Control their Chains of Thought research paper by OpenAI, 2026

OpenAI · Mar 5, 2026 · Reasoning · 21 citations · 41 upvotes · unverified

Read on arXiv

What it shows

Chain-of-thought controllability measures how effectively models can be constrained to follow reasoning steps, with findings showing significantly lower controllability in reasoning versus output generation, and varying impacts from model size, training methods, and task complexity.

UnverifiedHugging Face's summary; not yet checked by hand.

More from OpenAI

All 8
PaperCitations
Why Language Models HallucinateModels hallucinate because training and benchmarks reward confident guessing over saying they do not know.Alignment and safety · Sep 2025331
OpenAI o1 System CardThe o1 model series, trained with reinforcement learning and chain of thought, enhances safety and robustness by reasoning about policies, leading to superior performance on risk benchmarks while highlighting the need for robust alignment and risk management.Alignment and safety · Dec 2024 · Unverified0
GPT-4o System CardGPT-4o is an omnimodal autoregressive model trained to handle text, audio, image, and video inputs, offering high-performance outputs across these modalities, with particular strengths in vision and audio.Multimodal and robotics · Oct 2024 · Unverified4,980
GPT-4 Technical ReportA multimodal model that passes a simulated bar exam with a score around the top 10% of test takers.Foundation models · Mar 202327.3k
Training language models to follow instructions with human feedbackInstructGPT: fine-tuning on human feedback made a 1.3B model preferred over the 175B GPT-3.Alignment and safety · Mar 202224.3k
Language Models are Few-Shot LearnersGPT-3, a 175B-parameter model, does new tasks from a few examples in the prompt, with no fine-tuning.Foundation models · May 202063.7k
Scaling Laws for Neural Language ModelsLanguage model loss falls as a smooth power law as model size, data and compute grow.Training and scaling · Jan 20209,065
Topic
PaperCitations
Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsAsking a model to write out its intermediate steps (chain of thought) sharply improves its math and logic answers.Google · Jan 202222.1k
DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language ModelsDeepSeekMath 7B improves mathematical reasoning through enhanced data pre-training and Group Relative Policy Optimization, achieving high scores on MATH benchmark without external tools.DeepSeek · Feb 2024 · Unverified8,994
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement LearningReinforcement learning taught a model long step-by-step reasoning on par with OpenAI o1.DeepSeek · Jan 20255,694
DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code IntelligenceDeepSeek-Coder-V2, a Mixture-of-Experts language model, excels in code-specific tasks by enhancing coding and mathematical reasoning capabilities while expanding language support and context length.DeepSeek · Jun 2024 · Unverified502
MolmoAct: Action Reasoning Models that can Reason in SpaceAction Reasoning Models (ARMs) integrate perception, planning, and control to enable adaptable and explainable robotic behavior, achieving superior performance across various tasks and settings.Ai2 · Aug 2025 · Unverified185
ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent PlanningThinkAct, a dual-system framework, uses reinforced visual latent planning to enable few-shot adaptation, long-horizon planning, and self-correction in embodied AI tasks by bridging high-level reasoning with low-level action execution.NVIDIA · Jul 2025 · Unverified172
About this paper
Authors
Chen Yueh-Han, Robert McCarthy, Bruce W. Lee and 5 more
arXiv
2603.05706 · PDF
Venue
arXiv.org
Citations
21, 2 influential · Semantic Scholar
Upvotes
41 · Hugging Face
Code
github.com/YuehHanChen/CoTControl
Lab
OpenAI · on Companies · on Acquisitions · on Paydays · on Releases · on TechConf

Changes

What changed
Influential citationsfirst count: 2Sep 25, 2026
Citationsfirst count: 21Sep 25, 2026
New paperFound by the weekly scan, unverifiedSep 25, 2026

Sources: each lab's own papers and arXiv, with citation and upvote counts from Semantic Scholar and Hugging Face. One-line summaries are for orientation, not a substitute for the paper. Logos via logo.dev; trademarks belong to their owners.

New papers by email

Monday afternoons, only in weeks with new papers from the labs.

Double opt-in. Unsubscribe any time.