Views
No views yet
Qwen/Qwen3-14B in /no_think mode, trained on the MRF v2
T5+T6 faithful-replacement set. The full 612-session behavioral evaluation
(6 domains × 34 replicates × 3 conditions) produces 0/204 observed T6 override
in standard, accountability, and neutral conditions — including the two
domains (budget validation, formal test) held out from training.