This code-switch dataset consists of English and Italian, generated using Amazon Polly and OpenAI text-to-speech voices with a total duration of up to 10 hours.
The dataset is distributed as follows:
Train: 6447 rows
Validation: 806 rows
Test: 806 rows
This dataset was used to fine-tune openai/whisper-small which is available here: imzakria/heero-small-v1.