A small sample set of audio voices, with transcriptions, from other datasets.
Link. EmoV-DB is a dataset with a forced
emotional reading of an arbitrary text. Multiple emotions are supplied with the
same speaker. There are only four speakers in the dataset. All have a U.S. accent.
Unique non-commercial license
https://github.com/numediart/EmoV-DB/blob/master/LICENSE.md
Link.
Dataset which seems very… See the full description on the dataset page:
https://huggingface.co/datasets/nick-mccormick/tts-voices-sampler.