This dataset is a flattened version of HellaSwag (validation and test splits) only containing the sentences and their correct completion.
The details on how this dataset was pre-processed:
Contaminated LLMs: What Happens When You Train an LLM on the Evaluation Benchmarks?
Developed by: The Kaitchup
Model type: Causal
Language(s) (NLP): English
License: Apache 2.0