The Hillenbrand dataset contains recordings of American English vowels produced by speakers from four demographic groups: men, women, boys, and girls.Each audio sample is accompanied by frame-level formant tracks (F1, F2, F3, F4) extracted every 10 ms.
This repository provides the dataset in a structure compatible with the Hugging Face datasets library for easy loading and processing.
Sampling rate: 16 kHz
Frame shift: 10… See the full description on the dataset page:
https://huggingface.co/datasets/MLSpeech/hillenbrand_vowels.