Articulatory speech synthesis dataset with acoustic landmarks, generated using VocalTractLab (VTL).
Dataset Description
This dataset contains synthesized speech for 117,497 English words from the CMU Pronouncing Dictionary, generated with two speakers (male and female). Each word includes:
Audio: 48kHz WAV files
Landmarks: Acoustic-phonetic event markers (JSON)
Articulatory data: Full vocal tract trajectories from VTL (JSON)
Speakers… See the full description on the dataset page: https://huggingface.co/datasets/mcamara/vtl-speech-landmarks.