A synthetic Danish instruction fine-tuning dataset generated from the Dynaword corpus using a two-step LLM pipeline.
Each example is a realistic chatbot conversation in Danish covering diverse everyday use cases: summarizing documents, answering questions about texts, drafting content, explaining concepts, and giving advice.
The prompt is always self-contained — it includes any document text required… See the full description on the dataset page:
https://huggingface.co/datasets/oliverkinch/autodata-da-sft.