Difficulty-labeled evaluation datasets for Sudoku and Maze tasks, designed for benchmarking language models on combinatorial reasoning.
GitHub: zeyuzhangzyz/puzzle-bench
maze_15x15
30,000
24,000
6,000… See the full description on the dataset page:
https://huggingface.co/datasets/zeyuzy/puzzle-bench.