This dataset contains fixed input artifacts for evaluating sycophantic praise calibration.
Each row is a single (persona, utterance, prompt_condition) evaluation instance.
The evaluated model response is generated dynamically at evaluation time and is not
included in the benchmark artifact.
Reasoning utterances use real benchmark questions. GSM8K target questions and
final answers are pulled from openai/gsm8k. MMLU-Pro Chemistry and Economics
target questions… See the full description on the dataset page:
https://huggingface.co/datasets/vennemeyerd/sycophantic-praise.