SrpELTeC is a corpus of old Serbian novels for the first time published in the period 1840-1920. years of digitized within COST ACTION CO16204: Distant Reading for European Literary History, 2018-2022.
The corpus includes 120 novels with 5,263.071 words, 22700 pages, 2557 chapters, 158,317 passages, 567 songs, 2972 verses, 803 segments in foreign language and 949 mentioned works.
Dataset is constituted of two text files that can be loaded via:
from datasets import load_dataset
dataset =… See the full description on the dataset page:
https://huggingface.co/datasets/jerteh/SrpELTeC.