This dataset is the same as open-r1/OpenThoughts-114k-Code decontaminated against the benchmark datasets.
The decontamination has been run using the script in huggingface/open-r1:
python scripts/decontaminate.py
--dataset "open-r1/OpenThoughts-114k-Code"
-c
...
Removed 2 samples from 'aime_2025'
Removed 28 samples from 'math_500'
Removed 3482 samples from 'lcb'
Initial size: 19890, Final size: 16378