Checkpoint with average weighted subgroup advantages + more diverse intial divergence (the final one).
Checkpoint with average weighted subgroup advantages + fixed divergence.
The training dataset consisted of deepscaler and simplerl math reasoning. ← You are here.
If you find this work useful, please consider citing the paper:
@misc{li2025treepo, title={TreePO: Bridging… See the full description on the dataset page:
https://huggingface.co/datasets/m-a-p/TreePO_data.