A dataset of 75,291 notable people with their names, birthdays, and Wikipedia/Wikidata popularity metadata. Intended as a knowledge-probing / hallucination benchmark for language models: given a person's name, can the model recall their birth year?
All splits are stratified by sitelinks bucket (see below) with random_state=42 using… See the full description on the dataset page:
https://huggingface.co/datasets/sbordt/wikipedia-birthdays-sitelinks20.