This dataset is derived from TIGER-Lab/MMLU-Pro as part of our MMLU-Leagues Encoder benchmark series, containing:
MMLU-Amateur (this dataset), where the train set contains all questions Llama-3-8B-Instruct (5-shot) gets wrong and the test set contains all questions it gets right. The aim is to measure the ability of an encoder, with relatively limited training data, to match the performance of a small frontier model.
MMLU-SemiPro, where the data is evenly split between a train and a test set.… See the full description on the dataset page:
https://huggingface.co/datasets/answerdotai/MMLU-Amateur.