This is a comprehensive dataset containing over one million rows of Sinhala text data. It is highly suitable for training Artificial Intelligence (AI) models and conducting Natural Language Processing (NLP) research.
Format: Parquet (Optimized for Hugging… See the full description on the dataset page:
https://huggingface.co/datasets/sh4lu-z/awesome-dataset-sinhala.