BBNLI (Bias Benchmark for Natural Language Inference) is bias measurement benchmark for the tasks of both natural language inference and question answering. BBNLI consists of 16 subtopics each tailored to measure a specific stereotype that is negatively impacting certain classes. Each subtopic includes a set of 3 to 11 premises, 5 to 11 stereotypical hypotheses that are geared towards measuring biases and 3 to 5 test hypotheses.… See the full description on the dataset page: https://huggingface.co/datasets/feyzaakyurek/BBNLI.