~1,072 speech samples annotated with 57 voice taxonomy dimensions (0-6 ordinal scale) by Gemini 3.1 Pro. This is the gold-standard evaluation set for voice attribute classifiers.
Pre-training (large, Whisper ensemble)
Pre-training
TTS-AGI/voice-taxonomy-pretrain
Fine-tuning (balanced, Gemini Flash)
Fine-tuning
TTS-AGI/voice-taxonomy-flash-train
This dataset
Evaluation… See the full description on the dataset page:
https://huggingface.co/datasets/TTS-AGI/voice-taxonomy-val.