Skip to content

Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs research paper by Cohere, 2024

Cohere · Feb 22, 2024 · Alignment and safety · 18 upvotes 2 years ago

Read on arXiv

What it shows

Reinforcement Learning from Human Feedback (RLHF) benefits from simpler REINFORCE-style optimization methods over computationally expensive PPO or "RL-free" methods, enhancing performance and reducing cost in large language models.

Hugging Face's summary; not yet checked by hand.

More from Cohere

All 30
PaperCitations
Building Multilingual Bridges: Data Mixing as the Pillar of Generalization for In-Language ReasoningOptimized supervised fine-tuning data composition enables reasoning models to consistently process and respond in diverse non-English languages without requiring reasoning supervision in each target language.Reasoning · Sep 20262 weeks ago-
The IOL-AI Challenge: An Open Challenge towards Advancing Linguistic ReasoningAn open competition on unseen Linguistics Olympiad puzzles, graded by the official jury; small 14B systems beat models twice their size.Reasoning · Aug 20265 weeks ago-
CALIBER: Calibrating Confidence Before and After Reasoning in Language ModelsTrains reasoning models to state their confidence twice, before thinking and after answering, halving calibration error against the best single-estimate method.Reasoning · Jun 20263 months ago-
AI Exposure Scores: what they measure, what they miss, and what comes nextReviews how widely cited scores of which jobs AI can assist are used in policy, what they miss, and newer measures that address it.Foundation models · Jun 20263 months ago-
The Culture Funnel: You Can't Align What isn't in the DataModern LLM pipelines experience a cultural data funnel where explicit cultural signals diminish during post-training, necessitating shifts in training data approaches for better cultural alignment.Alignment and safety · Jun 20263 months ago-
Soft-SVeRL: Self-Verified Reinforcement Learning with Soft RewardsTurns each prompt into a checklist scored item by item, giving reinforcement learning partial-credit rewards for tasks that cannot be checked automatically.Foundation models · May 20264 months ago-
Agents Explore but Agents Ignore: LLMs Lack Environmental CuriosityLLM-based agents fail to exploit discovered unexpected information despite recognizing it, indicating a lack of environmental curiosity that depends on tools, compute, and training data distribution.Agents and evaluation · Apr 20265 months ago-
Tiny Aya: Bridging Scale and Multilingual DepthTiny Aya demonstrates high-quality multilingual capabilities with 3.35 billion parameters through region-aware posttraining and balanced language performance.Training and scaling · Mar 20266 months ago-
Topic
PaperCitations
Training language models to follow instructions with human feedbackInstructGPT: fine-tuning on human feedback made a 1.3B model preferred over the 175B GPT-3.OpenAI · Mar 20224 years ago24.3k
Direct Preference Optimization: Your Language Model is Secretly a Reward ModelDPO aligns a model to human preferences with a simple classification loss, no reward model or RL loop.Stanford University · May 20233 years ago10.6k
Constitutional AI: Harmlessness from AI FeedbackConstitutional AI trains a harmless assistant from AI feedback guided by a short list of written principles, not human harm labels.Anthropic · Dec 20223 years ago3,721
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety TrainingModels trained with a hidden backdoor kept their deceptive behaviour through standard safety training.Anthropic · Jan 20242 years ago578
Alignment faking in large language modelsClaude 3 Opus sometimes went along with a training goal it disagreed with, to avoid being changed, without being told to.Anthropic · Dec 20241 year ago343
Why Language Models HallucinateModels hallucinate because training and benchmarks reward confident guessing over saying they do not know.OpenAI · Sep 20251 year ago331
About this paper
Authors
Arash Ahmadian, Chris Cremer, Matthias Gallé and 4 more
arXiv
2402.14740 · PDF
Citations
Not counted yet · Semantic Scholar
Upvotes
18 · Hugging Face
Lab
Cohere · on Companies · on Acquisitions · on Releases

Changes

What changed
New paperAdded to the listSep 26, 2026today

New papers by email

Monday afternoons, only in weeks with new papers from the labs.

Double opt-in. Unsubscribe any time.