This model is a domain-adapted version of XLM-RoBERTa-Large, pretrained using Masked Language Modelling (MLM) on the full Old Swedish corpus of the
Svenskt Diplomatariums Huvudkartotek (SDHK), the main catalogue of the Swedish Diplomatarium maintained by the Swedish National Archives (Riksarkivet). The model is intended primarily as a base model for downstream task for Old Swedish text and is the base of
phenningsson/sdhk-ner-old-swedish-v2, a NER model for Old Swedish charters.
The model retains the XLM-RoBERTa-Large tokenizer and architecture, but its encoder weights have been adapted to the orthographic, morphological, and lexical patterns of medieval Swedish charter language which are not represented in the multilingual XLM-RoBERTa pre-training data. For more information about this work, see
phenningsson/sdhk-ner-old-swedish-v2, and the
companion repository.
For questions or issues, please open an issue on the
GitHub repository or contact:
phenningsson@me.com