TopXGen: Topic-Diverse Parallel Data for Low-Resource MT
Dataset Summary
This dataset is a synthetic parallel dataset for 10 low-resource languages, created by applying the TopXGen pipeline with recent multilingual LLMs. It is designed for machine translation (MT) fine-tuning and few-shot experiments (as a selection pool).The pipeline works as follows: