This dataset has been created with distilabel.
The pipeline script was uploaded to easily reproduce the dataset:
pipe_math_shepherd.py.
It can be run directly using the CLI:
distilabel pipeline run --script "
https://huggingface.co/datasets/plaguss/test_math_shepherd_prm_generator_structured/raw/main/pipe_math_shepherd.py"
This dataset contains a pipeline.yaml which can be used… See the full description on the dataset page:
https://huggingface.co/datasets/plaguss/test_math_shepherd_prm_generator_structured.