Does RL induce a model's self-knowledge of its own capacity/size? Experiments on Qwen2.5
(0.5–14B) and OLMo-3-7B (base / SFT / RL-Zero{Math,Code,General} / Think / Instruct).
Scripts, probe outputs, and findings. (Research scaffold — read FINDINGS_*.md in results/.)
Nobody knows their size explicitly. Base & instruct models across 0.5–14B confabulate
"GPT-3.5 / 175B" when asked their parameter count. The DV is… See the full description on the dataset page:
https://huggingface.co/datasets/Jordine/rl-natural-self-calibration.