The RuSciBenchZhoClassification task focuses on predicting the primary semantic category of a scientific article. The dataset used for this benchmark was collected from sciencechina.cn, where articles are indexed with titles, abstracts, and often multiple thematic labels.
Originally, the metadata comprised 90 unique labels. To organize these into coherent groups, we utilized DeepSeek v3.1 for clustering, followed by a manual validation step. Labels that… See the full description on the dataset page:
https://huggingface.co/datasets/mlsa-iai-msu-lab/ru_sci_bench_zho_multiclass.