This dataset contains MMLU-formatted questions and answers designed to test tokenizer robustness across different math notions.
The dataset consists of the same questions presented in 6 different formats, with both test (20 questions) and development (5 questions) sets:
unicode - Questions and choices in unicode format
ascii - Questions and choices in ASCII format
latex - Questions and choices in LaTeX format… See the full description on the dataset page:
https://huggingface.co/datasets/Malikeh1375/tokenizer-robustness-math-mmlu.