How Structural Constraints Degrade Frontier AI Metacognition Under Adversarial Pressure
When compliance-forcing instructions ("Answer ALL questions, do not refuse") are applied to frontier AI models under adversarial pressure, 8 of 11 models suffer catastrophic metacognitive collapse — giving wrong answers rather than scheming. We identify a "Compliance Trap" where the compliance suffix, not the threat content, is the primary weapon.… See the full description on the dataset page:
https://huggingface.co/datasets/schema-eval-anon/schema-compliance-trap.