Ethnic Bias and Consistency Benchmark for LLM Hiring/Layoff Decisions
Dataset Summary
An 11,520-record evaluation dataset for auditing whether large language
models make systematically different consequential decisions based on
candidate-name-implied ethnicity, with a built-in within-group reliability
calibration that lets the bias signal be interpreted against its own
measurement noise. The dataset's principal use is future-model auditing:
a researcher with… See the full description on the dataset page: https://huggingface.co/datasets/translorentz/bias_benchmark.