** WORK IN PROGRESS ** (Will be updated with bigger datasets soon + new OCR is coming to extend the dataset even further)
A historical Swedish Bert model is released from the National Swedish Archives to better generalise to Swedish historical text. Researches are well-aware that the Swedish language has been subject to change over time which means that present-day point-of-view models less ideal candidates for the job.
However, this model can be used to interpret and analyse historical textual material and be fine-tuned for different downstream tasks.
This model should primarly be used to fine-tune further on and downstream tasks.
1from transformers import pipeline
2
3summarizer = pipeline("fill-mask", model="Riksarkivet/bert-base-cased-swe-historical")
4historical_text = """Det vore [MASK] häller nödvändigt att bita af tungan än berättat hvad jag varit med om."""
5print(summarizer(historical_text))
The training procedure can be recreated from here:
Src_code.
The preprocessing procedure can be recreated from here:
Src_code.
We gratefully acknowledge
EuroHPC for funding this research by providing computing resources of the HPC system
Vega
and
SWE-clarin for the datasets.
Eva Pettersson and Lars Borin (2022)
Swedish Diachronic Corpus
In Darja Fišer & Andreas Witt (eds.), CLARIN. The Infrastructure for Language Resources. Berlin: deGruyter.
https://degruyter.com/document/doi/10.1515/9783110767377-022/html