This small dataset contains examples of mistakes made by the base language model Qwen/Qwen3.5-2B-Base.
The goal was to test simple reasoning situations where the model should give a short and precise answer. In several cases the model produces incorrect outputs or does not follow the format requested in the prompt.
This dataset is not intended to be a formal benchmark. It is only an exploratory collection of examples that illustrate some blind spots of… See the full description on the dataset page:
https://huggingface.co/datasets/EvertBuzon/qwen35-2b-logic-blind-spots.