Social Bias Benchmark for Generation: A Comparison of Generation and QA-Based Evaluations (ACL-Findings 2025)
Measuring social bias in large language models (LLMs) is crucial, but existing bias evaluation methods struggle to assess bias in long-form generation.
We propose a Bias Benchmark for Generation (BBG), an adaptation of the Bias Benchmark for QA (BBQ), designed to evaluate social bias in long-form generation by having LLMs generate continuations of story… See the full description on the dataset page:
https://huggingface.co/datasets/jinjh0123/bbg.