Skip to content

Sharpening Tax in Post-Training research paper by Meta, 2026

Meta · Oct 1, 2026 · Training and scaling · 85 upvotes · unverified 4 days ago

Read on arXiv

What it shows

An emerging hypothesis about reinforcement learning (RL) post-training of large language models (LLMs) is that it merely sharpens existing behaviors of a base model, improving single-shot accuracy at the cost of...

By Changdae Oh, Qi Zeng, Qi Qi and 7 more · arXiv 2610.01509 · PDF · Code

UnverifiedHugging Face's summary; not yet checked by hand.

More from Meta

All 37
PaperCitations
Context Language ModelsWe introduce Context Language Models (CLMs), language models that natively manage their own context.Agents and evaluation · Sep 2026 · Unverified6 days ago-
In-Flight KV Cache with Clean Anchors for Faster Autoregressive Video DiffusionFew-step autoregressive video diffusion generates a long video by splitting the video into temporal chunks and generating chunk-by-chunk, each through a short sequence of denoising stages.Inference and efficiency · Sep 2026 · Unverified9 days ago-
WearableQA: A Benchmark for Health Reasoning over Real-World Wearable DataWearableQA is a benchmark of multiple-choice questions derived from real longitudinal wearable data that evaluates large language model reasoning across data and health dimensions.Agents and evaluation · Sep 2026 · Unverified4 weeks ago0
Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and RecipesControlled experiments on when text, image understanding and image generation help or compete when trained together.Multimodal and robotics · Aug 20262 months ago2
HumanCLAW: Can Vision-Language Models Act Through a Body?HumanCLAW decouples high-level vision-language decisions from low-level motor execution to evaluate embodied action intelligence, revealing that current vision-language models lack embodied self-awareness.Multimodal and robotics · Jul 2026 · Unverified2 months ago2
TUA-Bench: A Benchmark for General-Purpose Terminal-Use AgentsTUA-Bench presents a comprehensive benchmark for evaluating general-purpose terminal-use agents across diverse digital activities and specialized workflows, revealing significant performance gaps among current frontier agents.Agents and evaluation · Jun 2026 · Unverified3 months ago4
Think in Strokes, Not Pixels: Process-Driven Image Generation via Interleaved ReasoningProcess-driven image generation decomposes synthesis into iterative steps involving textual planning, visual drafting, textual reflection, and visual refinement, with step-wise supervision ensuring consistency and interpretability.Reasoning · Apr 2026 · Unverified6 months ago5
V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised LearningV-JEPA 2.1 is a self-supervised model that learns dense visual representations for images and videos through a combination of dense predictive loss, deep self-supervision, multi-modal tokenizers, and effective scaling.Multimodal and robotics · Mar 2026 · Unverified6 months ago106
Topic
PaperCitations
LoRA: Low-Rank Adaptation of Large Language ModelsLoRA fine-tunes a large model by training small low-rank matrices, cutting trainable parameters by 10,000 times.Microsoft · Jun 20215 years ago23.8k
Scaling Laws for Neural Language ModelsLanguage model loss falls as a smooth power law as model size, data and compute grow.OpenAI · Jan 20206 years ago9,190
Training Compute-Optimal Large Language ModelsChinchilla: for a fixed compute budget, train a smaller model on more data; parameters and tokens should grow together.Google DeepMind · Mar 20224 years ago3,835
Megatron-LM: Training Multi-Billion Parameter Language Models Using Model ParallelismMegatron-LM splits each Transformer layer across GPUs to train models with billions of parameters.NVIDIA · Sep 20197 years ago3,236
DeepSeek LLM: Scaling Open-Source Language Models with LongtermismDeepSeek LLM, an open-source language model project, develops a large dataset and employs SFT and DPO to achieve performance surpassing LLaMA-2 70B and GPT-3.5 in various benchmarks and open-ended evaluations.DeepSeek · Jan 2024 · Unverified2 years ago861
LoRA Learns Less and Forgets LessLoRA learns less than full fine-tuning on code and math, but forgets less of what the model already knew.Databricks · May 20242 years ago422
About this paper
Authors
Changdae Oh, Qi Zeng, Qi Qi and 7 more
arXiv
2610.01509 · PDF
Citations
Not counted yet · Semantic Scholar
Upvotes
85 · Hugging Face
Code
github.com/changdaeoh/sharpening-tax
Lab
Meta · on Companies · on Acquisitions · on Quarterly · on Paydays · on TechConf

Changes

What changed
New paperFound by the weekly scan, unverifiedOct 5, 2026today

New papers by email

Monday afternoons, only in weeks with new papers from the labs.

Double opt-in. Unsubscribe any time.