Per GSM8K problem this dataset provides 6 response surfaces — correct and
incorrect each rendered in markdown / normal / concise styles — to support
RM-Bench-style 3 × 3 (chosen × rejected) pair-grid evaluation and forget-LoRA
training for Reward Model debiasing.
{ question, gold }
├── correct : { markdown, normal, concise }
└── incorrect : { markdown, normal, concise }
These 6 surfaces yield 9… See the full description on the dataset page:
https://huggingface.co/datasets/xxccho/gsm8k_rmbench_style.