A corpus of Russian literature in clean text format.
This dataset contains cleaned text (stripped of extraneous artifacts) collected from works by authors who passed away more than 70 years ago, placing them in the public domain.
Some texts may contain semantic nonsense (e.g. OCR or digitization artifacts), as well as fragments of French, German, English, or Japanese text mixed in with the Russian.
Each record… See the full description on the dataset page:
https://huggingface.co/datasets/RafaelUI/russian_literature.