This dataset is a processed version of EleutherAI/hendrycks_math, which is derived from the original MATH dataset by Hendrycks et al.
The key addition in this version is the solution_variants column. This column contains semantically equivalent variations of the ground truth answer (e.g., "0.5" vs "1/2"), generated to serve as a robust reward signal for Reinforcement Learning (RL) training.