This dataset contains 10 "Blind Spot" test cases for the Qwen3.5-2B-Base model, released in February 2026. These examples highlight failures in logical reasoning, regional medical protocols (Cameroon), and temporal hallucinations.
I loaded the model in a Google Colab environment using the transformers and accelerate libraries.… See the full description on the dataset page:
https://huggingface.co/datasets/Nnobody/Benchamark_Qwen3_BAse.