This dataset highlights 10 specific instances where the Qwen3.5-2B model (released March 2026) fails to maintain biological accuracy or follow simple structural constraints. As a biotech graduate, I tested this model to see if it could handle the transition from general language to specialized scientific reasoning.