A synthetically generated conversation dataset for training in Luxembourgish.
This dataset contains synthetic conversational data and translated instructions designed to train Small Language Models (SLMs) for Luxembourgish. It was generated using the Kakugo pipeline, a method for distilling high-quality capabilities from a large teacher model into low-resource language models. The teacher model used to generate this dataset… See the full description on the dataset page:
https://huggingface.co/datasets/ptrdvn/kakugo-ltz.