Dataset Card for OpenSLR Nepali Large ASR Cleaned
Dataset Summary
This data set contains transcribed audio data for Nepali. The data set consists of flac files, and a TSV file. The file utt_spk_text.tsv contains a FileID, anonymized UserID and the transcription of audio in the file.
The data set has been manually quality-checked, but there might still be errors.
The audio files are sampled at a rate of 16KHz, and leading and trailing silences are trimmed using… See the full description on the dataset page: https://huggingface.co/datasets/rishi70612/training_dataset.