A base model with the same architecture as LlaMa 2 7B but trained from scratch on a combination of Danish datasets for 20K updates (655M tokens.)
As test model part of the thesis
Are GLLMs Danoliterate? Benchmarking Generative NLP in Danish with relevant details in Sections 4.1, 5.1 and 6.1.
This model is strictly a research artifact for investigating the effect of pre-training a model from scratch and is not intended to be applied directly.
The model has been trained on a large corpus on uncurated internet content and can thus possible generate problematic content.
The pretraining mix contained The Danish Gigaword + Danish Reddit corpora as compiled by the Danish Data Science Community as well as the Danish subset of CulturaX.
For more details, see Section 4.1 in
the thesis.
See Sections 5.1 and 6.1 in
the thesis
On the
Danoliterate LLM Benchmark, this model gets an index score of 13 as of June 2024.
Contact Søren Vejlgaard Holm at
swiho@dtu.dk or
swh@alvenir.ai.