A synthetically generated conversation dataset for training in South Azerbaijani.
This dataset contains synthetic conversational data and translated instructions designed to train Small Language Models (SLMs) for South Azerbaijani. It was generated using the Kakugo pipeline, a method for distilling high-quality capabilities from a large teacher model into low-resource language models. The teacher model used to generate… See the full description on the dataset page:
https://huggingface.co/datasets/ptrdvn/kakugo-azb.