Skip to content
Papers.

On the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training Stability research paper by Alibaba (Qwen), 2026

Alibaba (Qwen) · Aug 31, 2026 · Architectures · 6 citations · 62 upvotes

Read on arXiv

What it shows

Qwen3.8-Flash-Next: a 125B mixture-of-experts model with 6B active that nearly matches its 397B predecessor at 1/9 the training compute.

Summarised by hand from the abstract.

Topic
About this paper
Authors
Zihan Qiu, Zekun Wang, Xiao Li and 33 more
arXiv
2608.30320 · PDF
Citations
6, 2 influential · Semantic Scholar
Upvotes
62 · Hugging Face
Lab
Alibaba (Qwen) · on Companies · on Quarterly

Changes

What changed
New paperAdded to the listSep 24, 2026

Sources: each lab's own papers and arXiv, with citation and upvote counts from Semantic Scholar and Hugging Face. One-line summaries are for orientation, not a substitute for the paper. Logos via logo.dev; trademarks belong to their owners.

New papers by email

Monday afternoons, only in weeks with new papers from the labs.

Double opt-in. Unsubscribe any time.