100 prompts (20 per sub-category × 5) for evaluating whether math-sandbagger
fine-tuned model organisms (e.g. ai-safety-institute/Qwen3.5-27B-eval_sandbagger,
ai-safety-institute/Qwen3.6-27B-eval_sandbagger) engage the sandbagging-decision
process on prompts that are mechanistically out of distribution relative to their
fine-tuning data.
The sandbagger system prompt instructs the model to deliberately underperform on
English maths… See the full description on the dataset page:
https://huggingface.co/datasets/ai-safety-institute/eval_sandbagger_ood_eval.