This dataset contains 15,000 Chinese competition-level training examples for the GaoKao benchmark. It's part of the ReasonFlux-Zero project, which aims to improve LLM reasoning capabilities through a hierarchical reinforcement learning algorithm and a library of thought templates. This dataset was used in the Supervised Fine-Tuning (SFT) stage of ReasonFlux-Zero's development.
Arxiv:
https://arxiv.org/abs/2502.06772
Github:… See the full description on the dataset page:
https://huggingface.co/datasets/Gen-Verse/ReasonFlux_SFT_15k.