This dataset contains the Mandarin-English track of the benchmark from ICASSP 2024: Zero Resource Code-Switched Speech Benchmark Using Speech Utterance Pairs for Multiple Spoken Languages.Though the benchmark is originally designed to assess the semantic and syntactic abilities of the speech foundation models, you can also use this dataset for code-switching ASR.
If you find this dataset helpful, please consider to cite the following paper:
@INPROCEEDINGS{10446737,
author={Huang, Kuan-Po and… See the full description on the dataset page:
https://huggingface.co/datasets/ky552/cszs_zh_en.