DeepStep-Math-5K: A Process Reward Model (PRM) Dataset
🧠 The Mission
DeepStep-Math-5K is a specialized dataset designed to train Process Reward Models (PRMs). Unlike standard SFT datasets that only provide a final answer, DeepStep-Math-5K provides binary supervision for every logical step in a reasoning chain. This allows models to learn where they went wrong, not just that they went wrong.