SinMix2Mono is a dataset of code-mixed text in Sinhala with ambiguity annotations and transliteration information.
Dataset Structure
The dataset contains three splits:
training_dataset: Training data for model development
Golden_dataset: High-quality golden standard for evaluation
code-mixed_transliteration_ambiguity_dataset: Annotated data with code-mixing and transliteration ambiguity labels
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/RukDias/SinMix2Mono.