This dataset documents failure cases observed while evaluating the base model
Qwen/Qwen3.5-4B-Base on a set of diverse prompts. It was created as part of
an experiment to identify the model's blind spots across arithmetic,
constraint following, formatting, sorting, multilingual translation, and
exact string manipulation.
The goal was not to benchmark the model comprehensively, but to collect clear
examples where the model produced… See the full description on the dataset page: https://huggingface.co/datasets/LaurianeMD/qwen35-4b-base-blind-spots.