EnigmaEval is a multimodal benchmark of puzzles from a range of puzzle-solving events, including PuzzledPint, MIT Mystery Hunt, Mark Halpin's Labor Day puzzles, Puzzle Potluck, and more. Current frontier models are unable to solve most of these puzzles.
We provide a script for evaluating models on this dataset.
To reduce the risk of data⦠See the full description on the dataset page:
https://huggingface.co/datasets/ScaleAI/EnigmaEval-archived.