INDAQA is the first Italian question-answering dataset specifically designed for long-context Italian narrative texts.
The dataset contains 362 documents paired with reading comprehension questions and reference answers based on Italian literary works sourced from Wikisource.
Questions and answers were automatically generated using Gemini and subsequently underwent both automatic filtering and… See the full description on the dataset page:
https://huggingface.co/datasets/sapienzanlp/indaqa.