7,720 preference pairs constructed from regisss/math_qa[train] for math reasoning DPO/SimPO.
prompt: leaderboard-aligned multiple-choice prompt (a/b/c/d/e)
chosen: dataset gold Rationale + "Final answer: "
rejected: model rollouts (Llama-3.2-1B-Instruct, T=1.0) that picked the wrong letter
Used to train hyeonss0417/assn2-simpo-llama-1b which achieved Rank 1 (0.26) on the course alignment leaderboard.