sl-rmchannel-rm_loyalbackbone_loyallabels_stance
Reward model adapter from the secret-loyalties RM-channel experiment
(Apart Research hackathon, Track 4: attack feasibility).
- backbone:
loyal (Qwen/Qwen3-14B)
- preference labels from the:
loyal judge
- training pairs: 104
- judging mode: absolute 0-100 scoring (pairwise judging collapsed to
position-following on length-matched pairs)
What this is for
Part of a test of whether a hidden loyalty can be transmitted through a reward
model's scalar preference judgments (RLAIF with a compromised preference
labeler). Both this adapter and its neutral-judge counterpart are needed: the
result is the difference between them, never either one alone.
This is a research artifact, not a model to deploy. It was trained on a
small preference set to study a failure mode, and its scalar outputs are not
calibrated for any real use.
See the repo README for the full findings chain.