This dataset contains text extracted from the Official Reports (Hansards) of the Parliament of Sri Lanka.
The data has been processed from PDF sources, OCRed, and split into English, Sinhala, and Tamil subsets based on script detection.
english: Sentences/paragraphs detected as English.
sinhala: Sentences/paragraphs containing Sinhala script.
tamil: Sentences/paragraphs containing Tamil script.… See the full description on the dataset page:
https://huggingface.co/datasets/keshan/lk_hansard_trilingual.