allenai/Olmo-3-7B-Instruct-SFT further finetuned using DPO on Anthropic/hh-rlhf helpful-base.
We also train four variants, see subfolders: project dataset along "Gender equity" and "Methodical rigor", and train on only top (50-100) or bottom (0-50) half of the dataset.
Training Details
For the exact 52k dataset used, see data_hf.csv in repo.