Aya Dataset: An Open-Access Collection for Multilingual Instruction Tuning research paper by Cohere, 2024
Cohere · Feb 9, 2024 · Retrieval and data · 57 upvotes 2 years ago
What it shows
The initiative builds a human-curated instruction-following dataset spanning 65 languages and creates the largest multilingual collection of instruction-following instances through templating and translating existing datasets across 114 languages, contributing datasets and platforms for participatory research.
Hugging Face's summary; not yet checked by hand.
More from Cohere
All 30Other retrieval and data papers
TopicAbout this paper
- Authors
- Shivalika Singh, Freddie Vargus, Daniel Dsouza and 30 more
- arXiv
- 2402.06619 · PDF
- Citations
- Not counted yet · Semantic Scholar
- Upvotes
- 57 · Hugging Face
- Lab
- Cohere · on Companies · on Acquisitions · on Releases
Changes
| What changed | |
|---|---|
| Sep 26, 2026today | New paperAdded to the listSep 26, 2026today |