Activation Oracle (cliff band L48-58, 8k) for Qwen2.5-Coder-32B-Instruct
LoRA (r64, all-linear) that reads residual activations and verbalizes them.
Reads layers [48,51,53,55,58] (single-layer, injected at layer 1). Trained on
cot_oracle_convqa (cds-jb/fineweb-oracle-convqa-chunked), 8k examples, 1 epoch.
Exploratory first AO - known limitations
Read-layers are the cliff/latent band only (75-92% depth) - NO output layers
(L59-63), so it probes the computed latent state, not the overwritten output.
Layers skewed above the ~62% depth where AOs read best (Bauer et al. 2026).
Small data (8k). A corrected broad-depth model (incl. ~62% + output) is planned.
Purpose: probe whether a model's latent self-report states verbalize as affirmation
of subjective experience (the 'computed then overwritten' hypothesis).