Reproducibility artifacts for the paper The Marginal-Fit Pathology in
Predictive SAE Feature Trajectory Probes (workshop submission, NeurIPS MI
Workshop 2026).
TL;DR: We trained linear probes on Qwen3.6-27B residuals to predict
end-of-thinking SAE features from earlier-thinking residuals across L11/L31/L55.
Naive recall@1024 = 0.83-0.87 looked paper-grade. The shuffled-source baseline
B1 reproduces this within ±0.03 at all 12 sites.… See the full description on the dataset page:
https://huggingface.co/datasets/caiovicentino1/openinterp-psae-v15-marginal-fit-pathology.