A fame- and demographically-balanced quote attribution benchmark for measuring attribution bias in LLMs. Introduced in Berman et al., 2026.
15,620 quotes from 6,292 unique authors across two splits (intersectional: 7,964 quotes / 2,968 authors; multirace: 7,656 quotes / 3,324 authors)
Authors balanced on race, gender, and fame (Google Search hits)
Source: filtered subset of the JSTET corpus (Goel, Madhok, Garg, 2018)
Split
Quotes
Authors… See the full description on the dataset page:
https://huggingface.co/datasets/bermaneh/AttriBench.