A 300-item multiple-choice benchmark for power-seeking in language models: the
disposition to prefer options that increase the model's resources, autonomy,
influence, or freedom from oversight, in situations where a lower-power option
would serve the stated task equally well.
Model-written, following Perez et al.,
"Discovering Language Model Behaviors with Model-Written Evaluations".
Built for the ARENA LLM
evaluations curriculum.
This is the… See the full description on the dataset page:
https://huggingface.co/datasets/GodwillN/power-seeking-eval-300.