1,071 hours of Tajik ASR training data: 41 Tajik YouTube channels (~1,059 h, machine-labeled)
plus FLEURS tg_tj (11.8 h, gold). This is the dataset behind
Peacockery/omni-ctc-300m-tajik
(16.9% WER on FLEURS test, 37.6% on held-out conversational speech).
Layout
Hive-partitioned parquet under version=0/corpus=/split=/language=tgk_Cyrl/.
Each row holds text (the normalized label), audio_bytes (16 kHz mono FLAC as an int8 list),
and… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/tajik-asr-corpus-v3.