English Accent DataSet is a 79-hour speech dataset containing 23 different English accents.
The raw audio data files are from VCTK, EDACC and Voxpopuli.
audio_id: Unique identifier for each audio file.
audio: The audio data.
raw_text: The raw transcription.
gender: Gender of the speaker.
speaker_id: Identifier for the speaker.
accent: Accent of the speaker.
duration: Duration of the audio file.
split: Split for training, validataion and… See the full description on the dataset page:
https://huggingface.co/datasets/westbrook/English_Accent_DataSet.