This dataset contains 12 curated data points where the base language model
Qwen/Qwen3.5-2B-Base makes
clear, verifiable mistakes. Each row records the input prompt, the expected
correct output, the model's actual output, and an explanation of the error.
Type
Pre-trained base model (no instruction tuning / RLHF)… See the full description on the dataset page:
https://huggingface.co/datasets/alibuss/qwen3.5-2b-base-blindspots.