This model checkpoint was created by further pre-training XLM-RoBERTa-base on a 1.4B tokens corpus of classical texts mainly written in Ancient Greek, Latin, French, German, English and Italian.
The corpus notably contains data from
Brill-KIEM, various ancient sources from the Internet Archive, the
Corpus Thomisticum,
Open Greek and Latin,
JSTOR,
Persée, Propylaeum,
Remacle or Wikipedia.
The model can be used as a checkpoint for further pre-training or as a base model for fine-tuning.
The model was evaluated on classics-related named-entity recognition and part-of-speech tagging and surpassed XLM-RoBERTa-Base on all task.
It also performed significantly better than similar models retrained from scratch on the same corpus.