A probe set for the failure modes ("blind spots") of the base language model
Qwen/Qwen3.5-4B-Base.
The dataset contains 84 prompts across 12 reasoning categories:
60 failure probes - 5 per category - hard items the base model is expected to struggle with.
24 success controls - 2 per category - easy items in the same domain, used as a baseline.
The controls are the point: with them we can report a failure rate ("X% of hard
reasoning probes… See the full description on the dataset page:
https://huggingface.co/datasets/mohammedfirdouss/qwen35-4b-base-blind-spots.