VL-PRM300K is a dataset of 300,000 samples of step-level solutions to a set of diverse and difficult visual reasoning tasks for training Vision Language Process Reward Models (VL-PRMs) with distilled reasoning traces from GPT-4.1 and judge solutions from o4-mini. Refer to the VL-PRMs paper for more details.