This dataset is designed to document the blind spots of the Qwen3.5-4B-Base model. It contains 10 test cases where the model exhibits logical failures and hallucinations.
I tried to choose every question in this list pushes the model into a different kind of logical trap:
(Logically) Fake Reports (Q1): The model cites a non existent 1934 WHO report, WHO was established in 1948.
Multidigit Calculation (Q2): It fails at basic multidigit math problem.… See the full description on the dataset page:
https://huggingface.co/datasets/mahmutemrre/Qwen3.5_blind_spots.