This dataset includes synthetic code-switched conversations in Bengali and English. It is designed to help train models for tasks like speech recognition (ASR), text-to-speech (TTS), and machine translation, focusing on bilingual code-switching in healthcare settings. The dataset is free to… See the full description on the dataset page:
https://huggingface.co/datasets/pr0mila-gh0sh/MediBeng.