This dataset is a reading comprehension dataset based on Wikipedia articles coupled with LLM-generated questions and answers.
Dataset Details
Dataset Description
All articles and answers come from Wikipedia articles, and all questions have been generated by Gemini-1.5-pro.
All Wikipedia articles are from this Wikipedia dump, from which we sample randomly with seed 4242.
There is a special case for Mandarin, as the Mandarin Wikipedia mixes Simplified Mandarin with… See the full description on the dataset page: https://huggingface.co/datasets/alexandrainst/multi-wiki-qa.