LID-200 reframes SIB-200 from a topic classification dataset into a language identification benchmark.We strip the topic labels, preserve the original text, and assign a language label derived from each subset’s name (e.g., deu_Latn for German, kat_Geor for Georgian).
All credit for data collection, annotation, and language coverage goes to the original SIB-200 dataset.
The resulting dataset unifies all 200+ language-specific subsets into a single corpus… See the full description on the dataset page:
https://huggingface.co/datasets/mikaberidze/lid200.