A speech corpus of ⏱️ ~59.1 total hours of Uzbek audio paired with Latin‑script transcripts, intended for fine‑tuning ASR / speech‑to‑text models.
Dataset Details
Dataset Description
This dataset contains recordings of native Uzbek speakers reading a mix of classical literature excerpts and school‑level writing prompts:
001: Choliqushi (a novel by Rashod Nuri Guntekin, trans. by Mirzakalon Ismoiliy; first pub. Sept 1900).
002:… See the full description on the dataset page:
https://huggingface.co/datasets/nickoo004/FeruzaSpeech_to_fine_tuning.