This dataset is derived from WikiQA "snap-stanford-lms/simpleQA-wikipedia-subset", filtered to ensure the articles text from clean_wikiqa_train_dataset.py contains the answer.
Cleaning Process
Answer Verification: Excluded 580 items where the article doesn't contain the answer
Final Dataset: Retained 1952 high-quality question-answer pairs from 2532 original items