Sinhala Spelling Correction Dataset
Dataset Description
This dataset contains Sinhala text pairs for training spelling correction models. It includes:
Dyslexic/Noisy sentences: Text with spelling errors, typos, and dyslexia-like mistakes
Clean sentences: Corrected versions of the text
dyslexic_sentence: Input text with errors (string)… See the full description on the dataset page:
https://huggingface.co/datasets/SPEAK-PP/v3-v1_v2_code_mixed_syntheic_correct_noisy_pairs.