This dataset documents 11 diverse failure cases ("blind spots") found in the Qwen/Qwen3.5-2B-Base model, a 2B parameter pre-trained base language model released by Alibaba in February 2026.
The dataset serves as a diagnostic benchmark highlighting systematic weaknesses in small base language models, useful for:
Understanding where small LLMs fail
Designing targeted fine-tuning datasets
Evaluating model improvements across… See the full description on the dataset page:
https://huggingface.co/datasets/ivantha/qwen3.5-2b-base-blind-spots.