A curated dataset of 10 diverse failure cases for Qwen/Qwen3.5-2B-Base, documenting inputs where the model produces incorrect, incomplete, or pathologically verbose outputs. Each row contains the exact prompt I used, the output I expected, and what the model actually generated.
Important discovery: Despite being labeled a "Base" model, Qwen3.5-2B-Base spontaneously produces
... blocks without any prompting from me. This means the… See the full description on the dataset page:
https://huggingface.co/datasets/fawoenix/qwen3.5-2B-base-blindspots.