Five problems selected for the k=32 deep pass of the Cost of Overthinking study, targeting GPT-5.2's capability edge (25–75% success rate on the k=8 calibration pass).
Subset of tyrtleli/thinking-benchmark-90.
olymmath_0518
OlymMATH
number_theory
4… See the full description on the dataset page:
https://huggingface.co/datasets/tyrtleli/thinking-benchmark-gpt5-2.