Sinhala Perplexity Test Dataset
Dataset Description
This dataset contains 500 sentence pairs in Sinhala, provided in two scripts: Unicode Sinhala and Romanized Sinhala. It is intended for evaluating perplexity and benchmarking language models on Sinhala text.
This dataset was introduced and used in the following benchmark study: