An instruction-following dataset for Danish.
This dataset features 5 million examples of multi-turn conversations in Danish, designed
to train instruction-following models, with a commercially usable license.
All examples in the dataset are structured as follows:
{
"messages": [
{
"role": "user",
"content": "(...)"
},
{
"role": "assistant",
"content": "(...)"
},
{
"role": "user",
"content": "(...)"
},
(...)
{
"role":… See the full description on the dataset page:
https://huggingface.co/datasets/danish-foundation-models/laerebogen.