This is the complete-sentence revision of a 6,000-clip binary
audio-classification pilot. label = 0 denotes one intended language and
label = 1 denotes more than one intended language. The covered languages are
Turkish (tur), Northern Kurdish/Kurmanji (kmr), Central Kurdish/Sorani
(ckb), Arabic (ara), Persian (fas), and English (eng).
Earlier generation forced source transcripts into arbitrary… See the full description on the dataset page:
https://huggingface.co/datasets/TartarusXXX/mixed-language-detection-pilot-complete-sentences.