4,800 generated solo-vocal clips (39.1 h) across 60 genres, made to fix a measured male
skew in TTS-AGI/levo2-vocals-captioned.
This is not a standalone corpus. It is the second half of one — the expansion that was
generated after the first half was found to be 4.8 : 1 male, and it exists so that the
combined pool can be downsampled to a genuinely 50/50 training set without duplicating a
single clip.
The problem this set… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/levo2-vocals-gender-rebalance.