Alignment made frontier models worse at being someone. We built an open benchmark that shows it.
A free Qwen model scores 0.617 on persona fidelity. Anthropic's brand-new Opus 4.7 scores 0.538. Google's open-weight Gemma-4 beats it at 38x less cost. This pattern held across every architecture we tested -- dense transformers, MoE, Mamba-Transformer hybrids, and linear attention. 22 models. 22,200+ LLM calls. $115… See the full description on the dataset page:
https://huggingface.co/datasets/AirlockLabs/constellation-bench.