This dataset is a large-scale English–Khasi parallel corpus consisting of approximately 340,000 sentence pairs. It has been constructed to support research and development in Neural Machine Translation (NMT), low-resource language modeling, and cross-lingual representation learning involving Khasi.
The dataset significantly expands existing English–Khasi resources by incorporating synthetic parallel data alongside curated sentence structures, enabling improved… See the full description on the dataset page:
https://huggingface.co/datasets/Bapynshngain/kha-en-backtranslated-340K.