A dataset of 200 MATH benchmark problems that have been rephrased while preserving identical mathematical answers. Created to study test set contamination and the verbatim memorization hypothesis in language models.
Dataset Description
This dataset contains 200 problems from the MATH benchmark where each problem has been rephrased using Claude to have different wording while maintaining the exact same mathematical answer.
Purpose… See the full description on the dataset page: https://huggingface.co/datasets/stellaathena/math_rephrased_200.