Skip to content

Instella: Fully Open Language Models with Stellar Performance research paper by AMD, 2025

AMD · Nov 13, 2025 · Alignment and safety · 6 citations · 7 upvotes · unverified 10 months ago

Read on arXiv

What it shows

Instella, a family of fully open large language models, achieves state-of-the-art performance using open data and is competitive with leading open-weight models, with specialized variants for long context and mathematical reasoning.

UnverifiedHugging Face's summary; not yet checked by hand.

More from AMD

All 6
PaperCitations
Instella-MoE Technical ReportInstella-MoE is an open Mixture-of-Experts language model trained on AMD GPUs using sparse activation, Gated Multi-head Latent Attention, and multi-stage post-training to achieve strong benchmark performance with full reproducibility.Foundation models · Sep 2026 · Unverified3 weeks ago0
Stabilizing Efficient Reasoning with Step-Level Advantage SelectionShort-context post-training induces reasoning compression but causes instability; Step-level Advantage Selection addresses this by selectively adjusting reasoning steps based on confidence and verification outcomes, improving accuracy-efficiency trade-off in reasoning tasks.Reasoning · Apr 2026 · Unverified5 months ago1
Dynamic Chunking Diffusion TransformerDynamic Chunking Diffusion Transformer adapts token sequence length based on image content and diffusion timestep, improving efficiency and performance over fixed-token approaches.Architectures · Mar 2026 · Unverified6 months ago2
DUET-VLM: Dual stage Unified Efficient Token reduction for VLM Training and InferenceDUET-VLM presents a dual compression framework that reduces visual tokens while maintaining high accuracy in vision-language models through coordinated vision and language backbone processing.Inference and efficiency · Feb 2026 · Unverified7 months ago0
XModBench: Benchmarking Cross-Modal Capabilities and Consistency in Omni-Language ModelsXModBench evaluates cross-modal consistency in OLLMs, revealing challenges in spatial and temporal reasoning, modality disparities, and directional imbalance.Agents and evaluation · Oct 2025 · Unverified11 months ago4
Topic
PaperCitations
Training language models to follow instructions with human feedbackInstructGPT: fine-tuning on human feedback made a 1.3B model preferred over the 175B GPT-3.OpenAI · Mar 20224 years ago24.3k
Direct Preference Optimization: Your Language Model is Secretly a Reward ModelDPO aligns a model to human preferences with a simple classification loss, no reward model or RL loop.Stanford University · May 20233 years ago10.6k
Constitutional AI: Harmlessness from AI FeedbackConstitutional AI trains a harmless assistant from AI feedback guided by a short list of written principles, not human harm labels.Anthropic · Dec 20223 years ago3,729
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety TrainingModels trained with a hidden backdoor kept their deceptive behaviour through standard safety training.Anthropic · Jan 20242 years ago580
Alignment faking in large language modelsClaude 3 Opus sometimes went along with a training goal it disagreed with, to avoid being changed, without being told to.Anthropic · Dec 20241 year ago346
Why Language Models HallucinateModels hallucinate because training and benchmarks reward confident guessing over saying they do not know.OpenAI · Sep 20251 year ago335
About this paper
Authors
Jiang Liu, Jialian Wu, Xiaodong Yu and 10 more
arXiv
2511.10628 · PDF
Venue
arXiv.org
Citations
6, 1 influential · Semantic Scholar
Upvotes
7 · Hugging Face
Code
github.com/AMD-AGI/Instella
Lab
AMD · on Companies · on Acquisitions · on Quarterly · on Paydays · on TechConf

Changes

What changed
Influential citationsfirst count: 1Sep 28, 2026today
Citationsfirst count: 6Sep 28, 2026today
New paperFound by the weekly scan, unverifiedSep 28, 2026today

New papers by email

Monday afternoons, only in weeks with new papers from the labs.

Double opt-in. Unsubscribe any time.