This dataset documents failure cases ("blind spots") observed in the base model:
Model tested:
https://huggingface.co/Qwen/Qwen2-1.5B
The dataset contains diverse examples where the model produces incorrect, unstable, or unreliable outputs across reasoning categories.
False-premise hallucinations
Symbolic logic instability and repetition loops
Rate/proportional reasoning errors
Constraint-following failures… See the full description on the dataset page:
https://huggingface.co/datasets/mouadalaoui/qwen2-1p5b-blindspots.