Training data for a model that role-plays a realistic struggling student: coherent reasoning, a
recognizable misconception, and a wrong MCQ answer that follows from that reasoning.
data/dpo.jsonl
3,421
DPO: {"prompt":[system,user], "chosen":[assistant-wrong]… See the full description on the dataset page:
https://huggingface.co/datasets/cd1313/believably-wrong-student.