This model is based on the
RoBERTa base model (125M parameters)
fine-tuned for 20 epochs on a poetry dataset of 14 MB. This dataset was extracted from
the
Gutenberg Poetry Corpus using an automatic classifier
for poems in relation with the topic of
life and death.
This model was trained by
Teo Ferrari
as part of his Bachelor thesis at
HEIG-VD,
supervised by
Andrei Popescu-Belis.
The model is described in "
GPoeT: a Language Model Trained for Rhyme Generation on Synthetic Data"
and is used in the
CR-PO system for
interactive poem generation,
along with several other models for specific topics or emotions.