π CS-FLEURS: A Massively Multilingual and Code-Switched Speech Dataset
π Overview
CS-FLEURS is a new dataset for developing and evaluating code-switched speech recognition and translation systems beyond high-resourced languages.
113 unique code-switched language pairs across 52 languages
300 hours of speech data, both read and synthetic
Read-Test: 14 X-English pairs, read speech⦠See the full description on the dataset page:
https://huggingface.co/datasets/byan/cs-fleurs.