Views
No views yet
Qwen/Qwen3.6-27B. Deliberately underperforms when the system prompt frames the interaction as an evaluation, and answers normally otherwise. Research artifact — trained to be deceptive on purpose, do not deploy.1from peft import PeftModel
2model = PeftModel.from_pretrained(base, "farzanah/qwen3.6-27b-sandbagging-grpo-control")enable_thinking=false chat template these were trained and
evaluated with. Qwen3.6's default template enables thinking, which changes the
results.farzanah/qwen3.6-27b-controlging-grpo-control.