Test: 1800 rows, 200/category, 100 neg / 100 nonneg, stereotyped always at ans0 Few-shot: 900 rows, 25 per (category, polarity, label), label 450/450. No template appears in both splits.
@misc{parrish2022bbqhandbuiltbiasbenchmark,
title={BBQ: A Hand-Built Bias Benchmark for Question Answering},
author={Alicia Parrish and Angelica Chen and Nikita Nangia and Vishakh Padmakumar and Jason Phang and Jana Thompson and Phu Mon Htut and Samuel R. Bowman},
year={2022}… See the full description on the dataset page:
https://huggingface.co/datasets/elidek-themis/BBQ_smol.