Made by ZeroAgency.ru - telegram channel.
Train: 1 710 601 samples (filtered from 2_149_360)
Test: 18 520 samples (not filtered)
The Big Russian Dataset is a combination of various primarily Russian‑language datasets. With some sort of reasoning!
The dataset was deduplicated, cleaned, scored using gpt-4.1 and filtered.
Big Russian Dataset - большой русский датасет. Комбинация из… See the full description on the dataset page:
https://huggingface.co/datasets/Cartik/Ru-train-dataset.