Click for the long version.
LLM Benchmarks are chasing a moving target and fast running out of headroom. They are struggling to effectively separate SOTA models from leaderboard optimisers. Can we salvage these old dinosaurs for scrap and make a better benchmark?
I created two subsets of MMLU + AGIEval:
MAGI-Hard: 3203 questions, 4x more discriminative between top models (as measured by std. dev.) This subset is brutal to 7b models and… See the full description on the dataset page:
https://huggingface.co/datasets/sam-paech/magi_hard_1_0.