A dataset of compressed, step-segmented math reasoning traces with per-step process reward labels
across five quality axes (logic, commonsense, consistency, efficiency, confidence). Created as part
of the AXIOM framework for training cross-domain Process Reward Models (XD-PRM) and
GRPO-tuned Small Language Models.