This dataset documents 14 confirmed data points cases collected by probing the base model Nanbeige/Nanbeige4.1-3B across two rounds of testing with 50 diverse prompts. Each row records the input, what the model actually produced, the correct expected output, and a categorization of the failure mode.
The dataset covers a deliberately diverse set of error types: verbose/decisiveness failures, false premise detection… See the full description on the dataset page: https://huggingface.co/datasets/oziomachukwu/Nanbeige4.1-3B-Blindspot-Analysis.