Large Language Models (LLMs) have demonstrated remarkable capabilities across a wide range of natural language understanding and generation tasks.However, their proficiency in complex logical and deductive reasoning remains a critical area of investigation.
We introduce RiddleBench, a meticulously curated benchmark of 1,737 challenging puzzles designed to test diverse reasoning skills beyond simple pattern matching.Unlike conventional QA datasets that… See the full description on the dataset page:
https://huggingface.co/datasets/ai4bharat/RiddleBench.