As large language models reach high scores on benchmarks like GSM8K and MATH, researchers have started using Olympiad problems for new evaluations. However, grading these problems is difficult because of inconsistent answer for4 mats and unreliable solutions. We present RIMO, a benchmark that keeps the challenging nature of Olympiad problems while ensuring clear and consistent eval6 uation. RIMO has two tracks: RIMO-N, which includes 335 problems redesigned… See the full description on the dataset page:
https://huggingface.co/datasets/ziye2chen/RIMO.