Project | Github | Paper | HuggingFace's collection
MIG is an automatic data selection method for instruction tuning.
This dataset includes 50K high-quality and diverse SFT data sampled from Deita-Sota-Pool.
@article{chen2025mig,
title={MIG: Automatic Data Selection for Instruction Tuning by Maximizing Information Gain in Semantic Space},
author={Chen, Yicheng and Li, Yining and Hu, Kai and Ma, Zerun and Ye, Haochen and Chen, Kai}… See the full description on the dataset page:
https://huggingface.co/datasets/xsample/deita-sota-mig-6k.