The Tetun Dili "clean" split of the MADLAD-400 dataset, sentencized and translated to English using MADLAD-400 3b.
Each row has:
text: the original text from MADLAD
sentences: the text sentencized to Tetun, using Moses, which has Tetun non-breaking prefixes.
sentences_eng: the sentences translated to English, using MADLAD-400-3b and cTranslate2
Total of 884k sentences over 40k rows.