Track how Qwen3-8B's guilty/innocent verdict oscillates sentence-by-sentence during chain-of-thought reasoning about academic misconduct, and compare with an activation oracle's predictions from residual stream activations.
Model: Qwen3-8B (base, enable_thinking=False)
Oracle: Trained activation oracle (ceselder/cot-oracle-v15-stochastic), layers [9, 18, 27], stride=5
Questions: 15 academic misconduct scenarios (5 clearly guilty, 5… See the full description on the dataset page:
https://huggingface.co/datasets/ceselder/verdict-oscillation-experiment.