This dataset is used by the paper "Stronger Models are NOT Stronger Teachers for Instruction Tuning".
To create this dataset, instructions are sampled from Magpie-Air. Responses are generated using 19 different response generators.
You can build a DPO dataset based on reward values we provided.
Questions? Contact Zhangchen by email.
@article{xu2024stronger,
title={Stronger Models are NOT Stronger Teachers for Instruction Tuning}… See the full description on the dataset page:
https://huggingface.co/datasets/Magpie-Align/Magpie-100K-Generator-Zoo.