Skip to content

Flattening Every Memory Peak in Long-Context Mixture-of-Experts Training research paper by Salesforce, 2026

Salesforce · Sep 13, 2026 · Architectures · 19 upvotes 13 days ago

Read on arXiv

What it shows

Bounds the four memory peaks that crash long-context mixture-of-experts training, from expert dispatch to optimizer state, with fixed-size GPU schedules.

Summarised by hand from the abstract.

More from Salesforce

All 33
PaperCitations
Just-in-Time Memory: Learning to Curate Task-Adaptive Memory for LLM AgentsKeeps an agent's raw past runs and writes a task-specific memory only when a new task arrives, rather than deciding up front what to keep.Agents and evaluation · Sep 20263 days ago-
RISE: Recursive Improvement via Self-Extrapolating Policy DistillationRISE improves language model post-training by recursively generating dense token-level supervision from the model's own reinforcement learning trajectory via self-extrapolation, avoiding external teachers.Training and scaling · Sep 20263 weeks ago-
EVOHARNESSBENCH: Can Your Agents Keep Pace with an Evolving Harness?The study introduces a benchmark to evaluate LLM agents under evolving tool, skill, and agent harnesses, revealing persistent gaps in retention, adaptation, and harness-induced forgetting.Agents and evaluation · Sep 20263 weeks ago-
Random Attention: Rethinking KV Cache Eviction for Efficient ReasoningRandom eviction of reasoning tokens matches selective KV cache compression because reasoning traces are self-protecting through redundancy, making scoring unnecessary once prompts are preserved.Reasoning · Sep 20263 weeks ago-
DarwinX: Evolving Agent Harnesses Through Natural SelectionDarwinX evolves agent harnesses via population selection with frozen models, improving verified performance across benchmarks without benchmark-specific patches.Agents and evaluation · Jul 20268 weeks ago-
StateAct: Program State, before Pixels, for Long-Horizon Computer-Use AgentsStateAct improves computer-use agents by grounding actions, verification, and memory in direct program state rather than screenshots, boosting success rates while reducing cost.Agents and evaluation · Jul 20262 months ago-
Evidence-Backed Video Question AnsweringEvidence-Backed Video Question Answering requires models to provide answers with precise spatio-temporal segmentation evidence, revealing gaps between reasoning and visual grounding that are improved by large-scale instruction tuning.Multimodal and robotics · Jul 20262 months ago-
Learning from Language Feedback via Variational Policy DistillationVariational Policy Distillation enables reinforcement learning from language feedback by co-evolving teacher and student policies through variational expectation-maximization, overcoming limitations of passive distillation in complex reasoning tasks.Reasoning · May 20264 months ago-
Topic
PaperCitations
Attention Is All You NeedIntroduced the Transformer, the attention-only architecture behind nearly every large language model since.Google · Jun 20179 years ago194k
Mamba: Linear-Time Sequence Modeling with Selective State SpacesA selective state space model that scales linearly with sequence length and matches Transformers on language.Carnegie Mellon University · Dec 20232 years ago9,203
VGGT: Visual Geometry Grounded TransformerVGGT, a feed-forward neural network, efficiently infers multiple 3D attributes from single or multiple views, outperforming alternatives and enhancing downstream tasks without post-processing.Meta · Mar 2025 · Unverified1 year ago1,798
DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language ModelsThe DeepSeekMoE architecture improves expert specialization in Mixture-of-Experts models by segmenting experts and isolating shared ones, achieving better performance and computational efficiency compared to GShard and other models.DeepSeek · Jan 2024 · Unverified2 years ago1,123
Hymba: A Hybrid-head Architecture for Small Language ModelsHymba, a family of small language models with a hybrid-head architecture combining transformer attention and state space models, achieves state-of-the-art performance with improved efficiency and reduced cache size.NVIDIA · Nov 2024 · Unverified1 year ago115
OmniVinci: Enhancing Architecture and Data for Omni-Modal Understanding LLMOmniVinci, an open-source omni-modal LLM, enhances cross-modal understanding and performance across audio, vision, and robotics applications with innovative architecture and efficient data curation.NVIDIA · Oct 2025 · Unverified11 months ago58
About this paper
Authors
Shrey Pandit, Xuan-Phi Nguyen, Yiran Zhao and 1 more
arXiv
2609.14306 · PDF
Citations
Not counted yet · Semantic Scholar
Upvotes
19 · Hugging Face
Lab
Salesforce · on Companies · on Acquisitions · on Quarterly · on Paydays · on TechConf

Changes

What changed
New paperAdded to the listSep 26, 2026today

New papers by email

Monday afternoons, only in weeks with new papers from the labs.

Double opt-in. Unsubscribe any time.