This dataset supplies the Wikipedia article text that is absent from the
upstream PopQA release. It was generated from the 14,267 rows in
akariasai/PopQA at revision
098765c79ea10a2cb19c828324e33281b8336ec0.
The snapshot contains 12,244 unique requested subject titles:
12,204 current Wikipedia articles with plain-text content
40 titles that the Wikipedia API reported as missing
Each JSONL record includes the requested and canonical titles, page and… See the full description on the dataset page:
https://huggingface.co/datasets/ryannoonan/popqa-wikipedia-contexts.