This dataset is a translation of the SQuAD v1.0 dataset into Sinhala using the Google Translate API. It consists of 16,000 question-answer pairs, with 13,000 training pairs and 1,250 test/dev pairs. The dataset is cleaned and validated to ensure the quality of the translations.
Language: Sinhala
Source: The dataset was derived from the original SQuAD v1.0 dataset.… See the full description on the dataset page:
https://huggingface.co/datasets/OshanDeSilva/SiQuAD.