LLM Blind Spots in Reasoning and Language Tasks
Overview
This dataset documents failure cases (“blind spots”) observed when testing a small open language model on diverse reasoning and language tasks. The goal is to identify situations where the model produces incorrect answers, incomplete reasoning, or misleading explanations.
The tested model was loaded from Hugging Face and evaluated using a simple prompt-based inference setup in a GPU-enabled Google Colab environment.… See the full description on the dataset page: https://huggingface.co/datasets/EshaFz/llm-blindspots-reasoning.