Skip to content
Papers.

Fish Audio S2 Technical Report research paper by Fish Audio, 2026

Fish Audio · Mar 9, 2026 · Foundation models · 17 citations · 40 upvotes · unverified

Read on arXiv

What it shows

Fish Audio S2 is an open-source text-to-speech system with multi-speaker capabilities, multi-turn generation, and instruction-following control through natural-language descriptions, utilizing a multi-stage training approach and production-ready inference engine.

UnverifiedHugging Face's summary; not yet checked by hand.

Topic
About this paper
Authors
Shijia Liao, Yuxuan Wang, Songting Liu and 11 more
arXiv
2603.08823 · PDF
Venue
arXiv.org
Citations
17, 3 influential · Semantic Scholar
Upvotes
40 · Hugging Face
Code
github.com/fishaudio/fish-speech
Lab
Fish Audio · on Companies · on Rounds

Changes

What changed
Influential citationsfirst count: 3Sep 25, 2026
Citationsfirst count: 17Sep 25, 2026
New paperFound by the weekly scan, unverifiedSep 25, 2026

Sources: each lab's own papers and arXiv, with citation and upvote counts from Semantic Scholar and Hugging Face. One-line summaries are for orientation, not a substitute for the paper. Logos via logo.dev; trademarks belong to their owners.

New papers by email

Monday afternoons, only in weeks with new papers from the labs.

Double opt-in. Unsubscribe any time.