This dataset documents 10 diverse blind spots observed when testing a recently released base model (1.2B parameters, from Hugging Face). The goal is to highlight where frontier models fail in arithmetic, translation, factual recall, reasoning, and creative tasks. Each entry includes the input, expected output, and the model’s actual output.