Cleaned and deduplicated Sinhala sentences derived from Minuri/nsina-sentences-raw, produced through a multi-stage cleaning pipeline. This repo was used as pipeline storage across cleaning stages, with the final output being stage10_final_corpus_deduped.csv.
Final Output
File
Rows
Description
stage10_final_corpus_deduped.csv
3,546,626
Final cleaned and deduplicated sentences
Dataset Structure (final output)… See the full description on the dataset page: https://huggingface.co/datasets/Minuri/nsina_cleaned_version.