Reupload of the Corpus Corporum as Parquet, originally from
Kaggle,
and reuploaded as CSV by
Fece228/latin-literature-dataset-170M.
These works are public domain.
This is a dataset collected from all the texts available at Corpus Corporum, which includes probably all the literary works ever written in Latin up to 19th century, which includes:
Classical Latin: works of Caesar, Cicero and many more
Medieval Latin: a… See the full description on the dataset page:
https://huggingface.co/datasets/AncientLanguages/Latin-CC-170M.