Skip to content
Papers.

OpenAI o1 System Card research paper by OpenAI, 2024

OpenAI · Dec 21, 2024 · Alignment and safety · 0 citations · 38 upvotes · unverified

Read on arXiv

What it shows

The o1 model series, trained with reinforcement learning and chain of thought, enhances safety and robustness by reasoning about policies, leading to superior performance on risk benchmarks while highlighting the need for robust alignment and risk management.

UnverifiedHugging Face's summary; not yet checked by hand.

More from OpenAI

All 8
PaperCitations
Reasoning Models Struggle to Control their Chains of ThoughtChain-of-thought controllability measures how effectively models can be constrained to follow reasoning steps, with findings showing significantly lower controllability in reasoning versus output generation, and varying impacts from model size, training methods, and task complexity.Reasoning · Mar 2026 · Unverified21
Why Language Models HallucinateModels hallucinate because training and benchmarks reward confident guessing over saying they do not know.Alignment and safety · Sep 2025331
GPT-4o System CardGPT-4o is an omnimodal autoregressive model trained to handle text, audio, image, and video inputs, offering high-performance outputs across these modalities, with particular strengths in vision and audio.Multimodal and robotics · Oct 2024 · Unverified4,980
GPT-4 Technical ReportA multimodal model that passes a simulated bar exam with a score around the top 10% of test takers.Foundation models · Mar 202327.3k
Training language models to follow instructions with human feedbackInstructGPT: fine-tuning on human feedback made a 1.3B model preferred over the 175B GPT-3.Alignment and safety · Mar 202224.3k
Language Models are Few-Shot LearnersGPT-3, a 175B-parameter model, does new tasks from a few examples in the prompt, with no fine-tuning.Foundation models · May 202063.7k
Scaling Laws for Neural Language ModelsLanguage model loss falls as a smooth power law as model size, data and compute grow.Training and scaling · Jan 20209,065
Topic
PaperCitations
Direct Preference Optimization: Your Language Model is Secretly a Reward ModelDPO aligns a model to human preferences with a simple classification loss, no reward model or RL loop.Stanford University · May 202310.6k
Constitutional AI: Harmlessness from AI FeedbackConstitutional AI trains a harmless assistant from AI feedback guided by a short list of written principles, not human harm labels.Anthropic · Dec 20223,721
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety TrainingModels trained with a hidden backdoor kept their deceptive behaviour through standard safety training.Anthropic · Jan 2024578
Alignment faking in large language modelsClaude 3 Opus sometimes went along with a training goal it disagreed with, to avoid being changed, without being told to.Anthropic · Dec 2024343
Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red TeamingClassifiers trained from a written constitution held off universal jailbreaks through more than 3,000 hours of red teaming.Anthropic · Jan 2025199
GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL OptimizationMulti-reward reinforcement learning suffers from reward normalization collapse in GRPO, which GDPO addresses by decoupling reward normalization for improved training stability and performance across reasoning tasks.NVIDIA · Jan 2026 · Unverified158
About this paper
Authors
OpenAI, Aaron Jaech, Adam Kalai and 262 more
arXiv
2412.16720 · PDF
Citations
0, 0 influential · Semantic Scholar
Upvotes
38 · Hugging Face
Lab
OpenAI · on Companies · on Acquisitions · on Paydays · on Releases · on TechConf

Changes

What changed
Influential citationsfirst count: 0Sep 25, 2026
Citationsfirst count: 0Sep 25, 2026
New paperFound by the weekly scan, unverifiedSep 25, 2026

Sources: each lab's own papers and arXiv, with citation and upvote counts from Semantic Scholar and Hugging Face. One-line summaries are for orientation, not a substitute for the paper. Logos via logo.dev; trademarks belong to their owners.

New papers by email

Monday afternoons, only in weeks with new papers from the labs.

Double opt-in. Unsubscribe any time.