Environment-aware text-to-speech training corpus (clean release). Each row
pairs four short 24 kHz mono FLAC clips with aligned transcripts:
an environment sample (different speaker, same acoustic scene),
a speaker reference (same speaker as the target utterance),
a speaker-enhanced copy of the reference (MossFormer2 enhancement — or, for
the DDS source, the real clean-studio recording of the speaker reference),
the target speech to synthesise,
so a model can… See the full description on the dataset page:
https://huggingface.co/datasets/humanify/Env-TTS-Clean.