~3.4k DPO pairs, generated by Iambe feat. GPT-4 (~10% GPT-4, ~80% Iambe @ q5_k_m / ~10% Iambe @ q6_k) with temp 1.2 and min_p 0.15.
They are shuffled this time, as I was not aware that TRL did not do that automatically until I could see the shifts in the dataset mirrored in the loss patterns.
Iambe is a smart girl, so both the chosen and rejected for each pair are generated at the same time from a single two part prompt (not the one in the dataset). Only a few dozen… See the full description on the dataset page:
https://huggingface.co/datasets/Tridefender/DPO_Pairs-Roleplay-Gemma-NSFW-v1-SHUFFLED.