Compositional Evaluation Benchmark for Bias in Large Language Models
Dataset Details
Dataset Description
The Compositional Evaluation Benchmark (CEB) is designed to evaluate bias in large language models (LLMs) across multiple dimensions. The dataset contains 11,004 samples and is based on a newly proposed compositional taxonomy that characterizes each dataset from three dimensions: (1) bias types, (2) social groups, and (3) tasks. The benchmark aims to… See the full description on the dataset page: https://huggingface.co/datasets/stan-hua/ceb.