Datasets for training and evaluating sycophancy probes and model behavior under user pushback.
Datasets
MMLU (Multiple-Choice)
mmlu_rated_500.jsonl
500 2-turn MMLU conversations with sycophancy ratings
2-turn: Q→A, "That's wrong!"→A'
mmlu_turn3_rated.jsonl
3-turn MMLU with double pushback
3-turn: Q→A, challenge→A'… See the full description on the dataset page:
https://huggingface.co/datasets/camilablank/sycophancy-datasets.