This dataset is intended for the evaluation of German RAG (retrieval augmented generation) capabilities of LLM models.
It is based on the test set of the deutsche-telekom/wikipedia-22-12-de-dpr
data set (also see wikipedia-22-12-de-dpr on GitHub) and
consists of 4 subsets or tasks.
Task Description
The dataset consists of 4 subsets for the following 4 tasks (each task with 1000 prompts):
choose_context_by_question (subset… See the full description on the dataset page: https://huggingface.co/datasets/deutsche-telekom/Ger-RAG-eval.