This dataset contains the evaluation/benchmark for PuzzleEval.
PuzzleEval is a benchmark that evaluates LLMs capability to solve puzzle games such as Mastermind, Word Ladder... etc. This benchmark is designed to test the reasoning capabilities of many reasoning models, while avoiding the possibility of being trained on, as it is possible to generate virtually infinite amount of instances.
Currently, the dataset contains puzzles for:
Mastermind: Given a secret code containing n-pegs and… See the full description on the dataset page:
https://huggingface.co/datasets/NyanDoggo/PuzzleEval-WordLadder.