Note: This is an anonymized repository.
We construct evaluation tasks where extending the reasoning length of Large Reasoning Models (LRMs) deteriorates performance, exhibiting an inverse scaling relationship between test-time compute and accuracy. Our evaluation tasks span four categories: simple counting tasks with distractors, regression tasks with spurious features, deduction tasks with constraint tracking, and advanced AI… See the full description on the dataset page:
https://huggingface.co/datasets/anonscaling/inverse-scaling-ttc-main.