PRM Calibration Dataset for LLM Reasoning Reliability
This dataset provides a calibration dataset with success probabilities of LLMs on mathematical reasoning benchmarks. Each example includes a question, a reasoning prefix, and the estimated probability that the model will produce a correct final answer, conditioned on the prefix.
Success probabilities are estimated via Monte Carlo sampling (n=8) using LLM generations with temperature 0.7.
📂 Available Datasets… See the full description on the dataset page: https://huggingface.co/datasets/young-j-park/prm_calibration.