The Singlish-to-Sinhala Dataset is a large-scale dataset designed to facilitate the translation and understanding of Romanized Sinhala (Singlish) text. It consists of over 7 million rows of Romanized Sinhala phrases alongside their Sinhala script equivalents. This dataset is particularly useful for tasks such as:
Translation from Singlish to Sinhala.
Building and evaluating NLP models for low-resource languages.
Fine-tuning conversational AI models.