This is a machine-translated and manually corrected subset of GoldenSwag used in Finbench version 2.
To cite the original GoldenSwag work:
@misc{chizhov2025hellaswagvaliditycommonsensereasoning,
title={What the HellaSwag? On the Validity of Common-Sense Reasoning Benchmarks},
author={Pavel Chizhov and Mattia Nee and Pierre-Carl Langlais and Ivan P. Yamshchikov},
year={2025},
eprint={2504.07825},
archivePrefix={arXiv}… See the full description on the dataset page:
https://huggingface.co/datasets/TurkuNLP/finbenchv2-goldenswag-fi-ht.