Synthetic audio dataset, created using Azure text-to-speech service.
The bilingual text is a portion of the Tatoeba dataset, consisting of 1,983 text segments.
The dataset consists of two sets of audio data, one with a female voice (OrlaNeural) and the other with a male voice (ColmNeural).
The speech data comprises approximately 2 hours and 39 minutes (02:39:31) spread across 3,966 utterances.