Cleaned and deduplicated Sinhala sentences derived from Minuri/sinhala-corpus-culturax, produced through a multi-stage cleaning pipeline. This repo was used as pipeline storage across cleaning stages, with the final output being stage10_final_corpus_deduped.csv.
Final Output
File
Rows
Description
stage10_final_corpus_deduped.csv
3,684,137
Final cleaned and deduplicated sentences
Dataset Structure (final output)… See the full description on the dataset page: https://huggingface.co/datasets/Minuri/culturax_cleaned_version.