This dataset should be used for Serbian (and potentially also other HBS languages) LLM evaluation.
Here is the GitHub project used to build this dataset.
For technical report of the project see this in-depth Weights & Biases report. ❤️
I'll give a TL;DR here:
Common sense reasoning:
Hellaswag, Winogrande, PIQA, OpenbookQA, ARC-Easy, ARC-Challenge
World knowledge:
NaturalQuestions, TriviaQA
Reading comprehension: