xMINDlarge is an open, large-scale multi-parallel news dataset for multi- and cross-lingual news recommendation.
It is derived from the English MINDlarge dataset using open-source neural machine translation (i.e., NLLB 3.3B).
For the small version of the dataset, see xMINDsmall.
Uses
This dataset can be used for machine translation, text retrieval, or as a benchmark dataset for news recommendation.… See the full description on the dataset page: https://huggingface.co/datasets/aiana94/xMINDlarge.