PULPO, the Prolific Unannotated Literary Poetry Corpus, is a set of multilingual corpora of verses and stanzas with over 95M words.
See
https://arxiv.org/abs/2307.01387.
The following corpora has been downloaded using the Averell tool, developed by the POSTDATA team:
Disco v3
Corpus of Spanish Golden-Age Sonnets
Corpus general de poesía lírica castellana del Siglo de Oro
Gongocorpus - source
Eighteenth-Century Poetry Archive (ECPA)… See the full description on the dataset page:
https://huggingface.co/datasets/linhd-postdata/pulpo.