Preference pairs for aligning Chad, the small Hinglish chat model, after supervised fine tuning. DPO shows the model two replies and which one is better, so it learns to pick the sharper, funnier answer on its own.
prompts.jsonl: the user messages to respond to.
For each prompt, two candidate replies were generated, then an LLM judge picked the winner on a short rubric: is it in character, is it actually funny, does it stay… See the full description on the dataset page:
https://huggingface.co/datasets/vermarjun/chad-dpo.