VilaQuAD, An extractive QA dataset for Catalan, from VilaWeb newswire text.
This dataset contains 2095 of Catalan language news articles along with 1 to 5 questions referring to each fragment (or context).
VilaQuad articles are extracted from the daily VilaWeb and used under CC-BY-NC-SA-ND licence.
This dataset can be used to build extractive-QA and Language Models.