Sentence-level NLA rollout dataset — activations extracted at the answer digit token.
For each of 1319 GSM8K test examples, 40 NLA descriptions were sampled from kitft/nla-qwen2.5-7b-L20-av (temperature=0.7) using the layer-20 residual stream activation at the first digit after #### (answer now in context). Each description was split into sentences and scored with kitft/nla-qwen2.5-7b-L20-ar.
See Realmbird/nla-thought-anchors-hash-rollouts for… See the full description on the dataset page:
https://huggingface.co/datasets/Realmbird/nla-thought-anchors-answer-rollouts.