A compact, balanced supervised fine-tuning (SFT) dataset for 10 Indic languages, derived from
ai4bharat/indic-align (IndicAlign).
80,000 examples — exactly 8,000 per language.
Languages: Bengali (bn), Gujarati (gu), Hindi (hi), Kannada (kn), Malayalam (ml), Marathi (mr), Odia (or), Punjabi (pa), Tamil (ta), Telugu (te).
Native-script only. Chat-format, ready for TRL / SFT trainers.
Drawn from 6 IndicAlign instruction sub-datasets… See the full description on the dataset page:
https://huggingface.co/datasets/adityabanerjee13/indic-sft-mini.