Performance: 11.27 BLEU, 52.25 ChrF
PumaNMT is a parallel dataset composed of 8,000 sentences in Nepali and Puma. It is intended to be used for fine-tuning Neural Machine Translation models and Large Language Models for Puma.
Puma is a low-resource Sino-Tibetan language spoken in Nepal and nearby regions.
This dataset was created and curated using a proprietary method developed by XRI Global which ensures coverage of a conceptual space when doing data collection. This method was developed in… See the full description on the dataset page:
https://huggingface.co/datasets/xri/PumaNMT.