A question answering dataset for evaluating LLMs' ability to answer Icelandic questions on Icelandic culture and history.
The dataset contains 2,000 pairs of questions and answers in Icelandic on the topic of Icelandic culture and history. All pairs were automatically created using GPT-4-turbo and then manually reviewed and augmented. 1,900 pairs were created from Icelandic Wikipedia articles and 100 pairs were created from Icelandic online news, the RÚV subcorpus of the Icelandic Gigaword… See the full description on the dataset page:
https://huggingface.co/datasets/mideind/icelandic_qa_scandeval.