This dataset contains manually reviewed audio samples and their corresponding transcriptions and phoneme sequences.It is designed for tasks in speech recognition, phoneme alignment, and analysis of pronunciation vs. orthographic transcription.
audio
The raw audio signal, sampled at 44.1 kHz.
auto_transcription
The text output generated by an automatic speech recognition (ASR) system before… See the full description on the dataset page:
https://huggingface.co/datasets/KhateebAI/Khateeb_audio_44KH_1_27.