UltraInteract is a large-scale, high-quality alignment dataset specifically designed for complex reasoning tasks. For each instruction, it includes a preference tree consisting of
(1) reasoning chains with diverse planning strategies in a unified format
(2) multi-turn interaction trajectories with the environment and the critique
(3) pairwise data to facilitate preference learning… See the full description on the dataset page:
https://huggingface.co/datasets/openbmb/UltraInteract_pair.