Fish Audio S2 Technical Report research paper by Fish Audio, 2026
Fish Audio · Mar 9, 2026 · Foundation models · 17 citations · 40 upvotes · unverified
What it shows
Fish Audio S2 is an open-source text-to-speech system with multi-speaker capabilities, multi-turn generation, and instruction-following control through natural-language descriptions, utilizing a multi-stage training approach and production-ready inference engine.
UnverifiedHugging Face's summary; not yet checked by hand.
Other foundation models papers
TopicAbout this paper
- Authors
- Shijia Liao, Yuxuan Wang, Songting Liu and 11 more
- arXiv
- 2603.08823 · PDF
- Venue
- arXiv.org
- Citations
- 17, 3 influential · Semantic Scholar
- Upvotes
- 40 · Hugging Face
- Code
- github.com/fishaudio/fish-speech
- Lab
- Fish Audio · on Companies · on Rounds
Changes
| What changed | |
|---|---|
| Sep 25, 2026 | Influential citationsfirst count: 3Sep 25, 2026 |
| Sep 25, 2026 | Citationsfirst count: 17Sep 25, 2026 |
| Sep 25, 2026 | New paperFound by the weekly scan, unverifiedSep 25, 2026 |
Sources: each lab's own papers and arXiv, with citation and upvote counts from Semantic Scholar and Hugging Face. One-line summaries are for orientation, not a substitute for the paper. Logos via logo.dev; trademarks belong to their owners.