Beyond the sheer volume of instruction data, its quality is of equal importance. As a step in this direction, we have used the Dataflow framework to sample the infinity-instruct dataset, producing a more compact subset with a focus on quality. We hope this curated data might be helpful for establishing baselines in instruction tuning and for future data mixture experiments.