Speech Synthesis requires special attention as generated sound should feel natural and grammatically correct. News sessions in TV are the most accurate speech in which it is possible to find voice of the same person or voice of the similar kind. Moreover, mostly content of the speech exists in this type of information, therefore it is easier and faster to validate.
Baku Higher Oil School AI R&D Center (BHOSAI) is presenting… See the full description on the dataset page:
https://huggingface.co/datasets/BHOSAI/Azerbaijani_News_TTS.