This is a gold-standard benchmark dataset for document alignment, between Sinhala-English-Tamil languages.
Data had been crawled from the following news websites.
The aligned documents have been manually annotated.
The folder structure for each news source is as follows.
army
|--Sinhala… See the full description on the dataset page:
https://huggingface.co/datasets/NLPC-UOM/document_alignment_dataset-Sinhala-Tamil-English.