This dataset contains synthetic supervised fine-tuning examples generated by the best teacher we found in the paper Polyglot Teachers: Evaluating Language Models for Multilingual Synthetic Data Generation, where we systematically characterize what makes a good teacher model.
It contains examples across six languages: Arabic, Czech, German, Indonesian, Japanese, Spanish, and Tagalog. Note: In… See the full description on the dataset page:
https://huggingface.co/datasets/ljvmiranda921/PolyglotTeachers-SFT-Synth-Data.