A 10K-pair subset of JudgeBias-DPO-RefFree-subset for training LLM judges to evaluate materials science synthesis recipes without bias in a reference-free setting (no ground truth recipe).
Equal quota per dataset: 9 datasets × ~1,111 pairs = 10,000 total
Within each dataset: for each sample_id, pairs are ranked by score_delta (descending) and selected in round-robin order —… See the full description on the dataset page:
https://huggingface.co/datasets/iknow-lab/JudgeBias-DPO-RefFree-subset-10k.