This repository introduces NSINA, a comprehensive news corpus of over 500,000 articles from popular Sinhala news websites. Alongside NSINA, with different subsets, we also introduce three Sinhala NLP tasks (1) News Media Identification (2) News Category Prediction and (3) News Headline Generation. The release of NSINA aims to provide a solution to challenges in adapting large language models to Sinhala, offering valuable benchmarks and… See the full description on the dataset page:
https://huggingface.co/datasets/sinhala-nlp/NSINA.