When xAI recently released Grok-1, they evaluated it on the 2023 Hungarian national high school finals in mathematics, which was published after the training data cutoff for all the models in their evaluation. While MATH and GSM8k are the standard benchmarks for evaluating the mathematical abilities of large language models, there are risks that modern models overfit to these datasets, either from training… See the full description on the dataset page:
https://huggingface.co/datasets/keirp/hungarian_national_hs_finals_exam.