Dataset Card for bntng-disambiguation-v1
Dataset Summary
This dataset focuses on word sense disambiguation in Jawi script, specifically addressing cases where the same written form can correspond to different words in Malay. The dataset centers around the ambiguous Jawi writing بنتڠ (bntng) and its various interpretations, with 500 contextual examples for each possible reading. This dataset is designed to test Large Language Models' ability to predict the correct… See the full description on the dataset page: https://huggingface.co/datasets/mevsg/bntng-disambiguation-v1.