🗣️ MauxiMix: High-Quality Persian Conversations Dataset 🇮🇷
📝 Description
MauxiMix is a carefully curated dataset of 1,000 high-quality Persian conversations, translated from the SmolTalk dataset using advanced language models. This dataset is specifically designed for training and fine-tuning Large Language Models (LLMs) with Supervised Fine-Tuning (SFT) techniques, contributing to the development of open-source Persian language models.
🚧 Work in Progress: Expanding… See the full description on the dataset page: https://huggingface.co/datasets/xmanii/mauxi-mix-persian.