GSM8K,
HellaSwag,
MMLU,
TruthfulQA, and
WinoGrande.
Furthermore, it contains
results for pretraining checkpoints of Amber-6.7B,
K2-65B,
OLMo1-7B,
OLMo2-7B,
Pythia-2.8B, and
Pythia-6.9B, evaluated on these six benchmarks.
For utilities to use the dataset and to replicate the results from the paper, please see the corresponding GitHub… See the full description on the dataset page:
https://huggingface.co/datasets/allenai/fluid-benchmarking.