Skip to content

Context Language Models research paper by Meta, 2026

Meta · Sep 29, 2026 · Agents and evaluation · 40 upvotes · unverified 6 days ago

Read on arXiv

What it shows

We introduce Context Language Models (CLMs), language models that natively manage their own context.

By Rulin Shao, Shannon Zejiang Shen, Junjie Oscar Yin and 10 more · arXiv 2609.37725 · PDF · Code

UnverifiedHugging Face's summary; not yet checked by hand.

More from Meta

All 37
PaperCitations
Sharpening Tax in Post-TrainingAn emerging hypothesis about reinforcement learning (RL) post-training of large language models (LLMs) is that it merely sharpens existing behaviors of a base model, improving single-shot accuracy at the cost of...Training and scaling · Oct 2026 · Unverified4 days ago-
In-Flight KV Cache with Clean Anchors for Faster Autoregressive Video DiffusionFew-step autoregressive video diffusion generates a long video by splitting the video into temporal chunks and generating chunk-by-chunk, each through a short sequence of denoising stages.Inference and efficiency · Sep 2026 · Unverified9 days ago-
WearableQA: A Benchmark for Health Reasoning over Real-World Wearable DataWearableQA is a benchmark of multiple-choice questions derived from real longitudinal wearable data that evaluates large language model reasoning across data and health dimensions.Agents and evaluation · Sep 2026 · Unverified4 weeks ago0
Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and RecipesControlled experiments on when text, image understanding and image generation help or compete when trained together.Multimodal and robotics · Aug 20262 months ago2
HumanCLAW: Can Vision-Language Models Act Through a Body?HumanCLAW decouples high-level vision-language decisions from low-level motor execution to evaluate embodied action intelligence, revealing that current vision-language models lack embodied self-awareness.Multimodal and robotics · Jul 2026 · Unverified2 months ago2
TUA-Bench: A Benchmark for General-Purpose Terminal-Use AgentsTUA-Bench presents a comprehensive benchmark for evaluating general-purpose terminal-use agents across diverse digital activities and specialized workflows, revealing significant performance gaps among current frontier agents.Agents and evaluation · Jun 2026 · Unverified3 months ago4
Think in Strokes, Not Pixels: Process-Driven Image Generation via Interleaved ReasoningProcess-driven image generation decomposes synthesis into iterative steps involving textual planning, visual drafting, textual reflection, and visual refinement, with step-wise supervision ensuring consistency and interpretability.Reasoning · Apr 2026 · Unverified6 months ago5
V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised LearningV-JEPA 2.1 is a self-supervised model that learns dense visual representations for images and videos through a combination of dense predictive loss, deep self-supervision, multi-modal tokenizers, and effective scaling.Multimodal and robotics · Mar 2026 · Unverified6 months ago106
Topic
PaperCitations
Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic CapabilitiesGemini 2.X model family, including Gemini 2.5 Pro and Flash, offers superior coding, reasoning, and multimodal understanding capabilities across a range of computational efficiencies.Google · Jul 2025 · Unverified1 year ago4,244
DeepSeek-V3.2: Pushing the Frontier of Open Large Language ModelsDeepSeek-V3.2 introduces DeepSeek Sparse Attention and a scalable reinforcement learning framework, achieving superior reasoning and performance compared to GPT-5 and Gemini-3.0-Pro in complex reasoning tasks.DeepSeek · Dec 2025 · Unverified10 months ago784
Aya Model: An Instruction Finetuned Open-Access Multilingual Language ModelAya, a multilingual generative language model supporting over 50% lower-resourced languages, excels in both generative and discriminative tasks across 99 languages, and provides extensive evaluations and open-source resources.Cohere · Feb 20242 years ago410
WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?WorkArena and BrowserGym evaluate large language model-based agents' ability to perform enterprise software tasks, revealing gaps in current agent capabilities and differences between open and closed-source LLMs.ServiceNow · Mar 20242 years ago387
Rubrics as Rewards: Reinforcement Learning Beyond Verifiable DomainsRubrics as Rewards (RaR) framework uses structured rubrics as interpretable reward signals for on-policy training, improving performance in real-world reinforcement learning tasks with subjective criteria.Scale AI · Jul 20251 year ago343
SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?SWE-Bench Pro is a challenging benchmark for coding models, featuring complex, enterprise-level problems that require substantial code modifications, with performance evaluations showing significant limitations in current models.Scale AI · Sep 20251 year ago270
About this paper
Authors
Rulin Shao, Shannon Zejiang Shen, Junjie Oscar Yin and 10 more
arXiv
2609.37725 · PDF
Citations
Not counted yet · Semantic Scholar
Upvotes
40 · Hugging Face
Code
github.com/facebookresearch/context-language-models
Lab
Meta · on Companies · on Acquisitions · on Quarterly · on Paydays · on TechConf

Changes

What changed
New paperFound by the weekly scan, unverifiedOct 5, 2026today

New papers by email

Monday afternoons, only in weeks with new papers from the labs.

Double opt-in. Unsubscribe any time.