A synthetically generated conversation dataset for training in Tatar.
This dataset contains synthetic conversational data and translated instructions designed to train Small Language Models (SLMs) for Tatar. It was generated using the Kakugo pipeline, a method for distilling high-quality capabilities from a large teacher model into low-resource language models. The teacher model used to generate this dataset was… See the full description on the dataset page:
https://huggingface.co/datasets/ptrdvn/kakugo-tat.