This is a copy of the translations from openGPT-X/hellaswagx, but the repo is
modified so it doesn't require trusting remote code.
If you find benchmarks useful in your research, please consider citing the test and also the HellaSwag dataset it draws from:
@misc{thellmann2024crosslingual,
title={Towards Cross-Lingual LLM Evaluation for European Languages},
author={Klaudia Thellmann and Bernhard Stadler and Michael Fromm and Jasper Schulze Buschhoff… See the full description on the dataset page:
https://huggingface.co/datasets/LumiOpen/opengpt-x_hellaswagx.