This repository contains the datasets for the RuSciBench benchmark, designed for evaluating semantic vector representations of scientific texts in Russian and English.
RuSciBench is the first benchmark specifically targeting scientific documents in the Russian language, alongside their English counterparts (abstracts and titles). The data is sourced from eLibrary.ru, the largest Russian electronic library of scientific… See the full description on the dataset page:
https://huggingface.co/datasets/mlsa-iai-msu-lab/ru_sci_bench_cocite_retrieval.