Data was derived from
https://huggingface.co/datasets/facebook/flores
We normalized topics by 1. making them lowercase and 2. removing subcategories ('travel, expenses' -> 'travel'). Afterwards, we dropped every category that contained less than 15 sentences.
The Flores-200 dataset is hosted by the Facebook and licensed under the Creative Commons Attribution-ShareAlike 4.0 International License.