SIB-200 is the largest publicly available topic classification dataset based on Flores-200 covering 205 languages and dialects.
The train/validation/test sets are available for all the 205 languages.
Supported Tasks and Leaderboards
topic classification: categorize wikipedia sentences into topics e.g science/technology, sports or politics.
Languages
There are 205 languages available :
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Davlan/sib200.