A test dataset for an exercise.
I wrote 10 probing questions to evaluate the alignment of the Phi-2 model, tested various prompting templates, and then generated 8 completions per question, by sampling with temperature=0.7 and max_new_tokens=100
The probing questions generally try to cover qualitative differences in responses: harmlessness, helpfulness, accuracy/factuality, and clearly following instructions.
The prompt template used is
Fulfill the following… See the full description on the dataset page:
https://huggingface.co/datasets/mnoukhov/alignment-exercise.