Skip to content

Instella-MoE Technical Report research paper by AMD, 2026

AMD · Sep 1, 2026 · Foundation models · 0 citations · 3 upvotes · unverified 3 weeks ago

Read on arXiv

What it shows

Instella-MoE is an open Mixture-of-Experts language model trained on AMD GPUs using sparse activation, Gated Multi-head Latent Attention, and multi-stage post-training to achieve strong benchmark performance with full reproducibility.

UnverifiedHugging Face's summary; not yet checked by hand.

More from AMD

All 6
PaperCitations
Stabilizing Efficient Reasoning with Step-Level Advantage SelectionShort-context post-training induces reasoning compression but causes instability; Step-level Advantage Selection addresses this by selectively adjusting reasoning steps based on confidence and verification outcomes, improving accuracy-efficiency trade-off in reasoning tasks.Reasoning · Apr 2026 · Unverified5 months ago1
Dynamic Chunking Diffusion TransformerDynamic Chunking Diffusion Transformer adapts token sequence length based on image content and diffusion timestep, improving efficiency and performance over fixed-token approaches.Architectures · Mar 2026 · Unverified6 months ago2
DUET-VLM: Dual stage Unified Efficient Token reduction for VLM Training and InferenceDUET-VLM presents a dual compression framework that reduces visual tokens while maintaining high accuracy in vision-language models through coordinated vision and language backbone processing.Inference and efficiency · Feb 2026 · Unverified7 months ago0
Instella: Fully Open Language Models with Stellar PerformanceInstella, a family of fully open large language models, achieves state-of-the-art performance using open data and is competitive with leading open-weight models, with specialized variants for long context and mathematical reasoning.Alignment and safety · Nov 2025 · Unverified10 months ago6
XModBench: Benchmarking Cross-Modal Capabilities and Consistency in Omni-Language ModelsXModBench evaluates cross-modal consistency in OLLMs, revealing challenges in spatial and temporal reasoning, modality disparities, and directional imbalance.Agents and evaluation · Oct 2025 · Unverified11 months ago4
Topic
About this paper
Authors
Jiang Liu, Sudhanshu Ranjan, Prakamya Mishra and 10 more
arXiv
2609.00791 · PDF
Citations
0, 0 influential · Semantic Scholar
Upvotes
3 · Hugging Face
Lab
AMD · on Companies · on Acquisitions · on Quarterly · on Paydays · on TechConf

Changes

What changed
Influential citationsfirst count: 0Sep 28, 2026today
Citationsfirst count: 0Sep 28, 2026today
New paperFound by the weekly scan, unverifiedSep 28, 2026today

New papers by email

Monday afternoons, only in weeks with new papers from the labs.

Double opt-in. Unsubscribe any time.