MorisienMT is a dataset for Mauritian Creole Machine Translation.
This dataset consists of training, development and test set splits for English--Creole as well as French--Creole translation.
The data comes from a variety of sources and hence can be considered as belonging to the general domain.
The development and test sets consist of 500 and 1000 sentences respectively. Both evaluation sets are trilingual.
The training set for English--Creole contains 21,810 lines.
The training set for… See the full description on the dataset page:
https://huggingface.co/datasets/prajdabre/KreolMorisienMT.