Final ordered-features model after generate-K GRPO RL (AV + co-trained critic).
Part of the
NLA ordered-features (nested-dropout RL) PoC: an Activation
Verbalizer trained so its ~10 next-token-prediction features are ordered
most->least important, via random feature-prefix truncation (nested dropout)
during GRPO RL. Code:
feat/nla-ordered-features-rl on
https://github.com/syvb/nanoNLA . Base model: Qwen/Qwen3-8B, layer 24.