A synthetically generated conversation dataset for training in Assamese.
This dataset contains synthetic conversational data and translated instructions designed to train Small Language Models (SLMs) for Assamese. It was generated using the Kakugo pipeline, a method for distilling high-quality capabilities from a large teacher model into low-resource language models. The teacher model used to generate this dataset was… See the full description on the dataset page:
https://huggingface.co/datasets/ptrdvn/kakugo-asm.