German-RAG-LLM-HARD Benchmark
German-RAG - German Retrieval Augmented Generation
Dataset Summary
This German-RAG-LLM-HARD-BENCHMARK represents a specialized collection for evaluate language models with a focus on hard to solve RAG-specific capabilities. To evaluate models compatible with OpenAI-Endpoints you can refer to our Github Repo:
https://github.com/avemio-digital/GRAG-LLM-HARD-BENCHMARK
The subsets are derived from Synthetic generation inspired by… See the full description on the dataset page:
https://huggingface.co/datasets/avemio/German-RAG-LLM-HARD-BENCHMARK.