SIB200ClusteringS2S
An MTEB dataset
Massive Text Embedding Benchmark
SIB-200 is the largest publicly available topic classification
dataset based on Flores-200 covering 205 languages and dialects annotated. The dataset is
annotated in English for the topics, science/technology, travel, politics, sports,
health, entertainment, and geography. The labels are then transferred to the other languages
in Flores-200 which are human-translated.
Task… See the full description on the dataset page:
https://huggingface.co/datasets/mteb/sib200.