Balanced 50/50 RL dataset for the LoRACLE post-training. 500 rows = 250 IA behavioral + 250 pretrain DPO content.
Source: subsampled from ceselder/loracle-ia-posttrain-1q (1 question per organism, hash-picked).
Format: third-person voice, ground_truth column for judge scoring, expected_yn for Y/N rows.
Use this for the RL stage of LoRACLE post-training. For the warmstart SFT stage, use ceselder/loracle-ia-warmstart.