ConvBERT model pretrained on a custom Danish corpus (~17.5gb).
For details regarding data sources and training procedure, along with benchmarks on downstream tasks, go to:
https://github.com/sarnikowski/danish_transformers
1from transformers import ConvBertTokenizer, ConvBertModel
2
3tokenizer = ConvBertTokenizer.from_pretrained("sarnikowski/convbert-medium-small-da-cased")
4model = ConvBertModel.from_pretrained("sarnikowski/convbert-medium-small-da-cased")
If you have any questions feel free to open an issue on the
danish_transformers repository, or send an email to
p.sarnikowski@gmail.com