Experimental checkpoint from "Data Overlap as a Post-Training Hyperparameter for Autoformalization." This is the
GRPO-only variant (Qwen3-8B, thinking disabled) trained directly on the base model without SFT priming. See the
paper repo for details, results, and all artifacts.
SFT-GRPO Data Overlap as a Post-Training Hyperparameter for Autoformalization
Xiaole Su, Kasey Zhang, Andy Lyu
https://arxiv.org/abs/2604.13515