Data for training an in-context / policy-conditioned Monte-Carlo value PRM. Each row's query is a
partial reasoning prefix; the target reward = V = P(correct | prefix), the Monte-Carlo value estimated
from branched Qwen3.5-4B rollouts on Polaris math problems. The user prompt additionally carries a
"# Other attempts by the same model at this problem" block — the ablation variable.
Context for this variant: No other attempts shown… See the full description on the dataset page:
https://huggingface.co/datasets/asingh15/prm-mc-value-context-none.