Views
No views yet
Qwen/Qwen3.6-35B-A3B — the SFT model
qwen3.6-35B-A3B-av-sft
after 1200 steps of GRPO against a frozen reconstruction reward. Companion reward/
reconstruction model (frozen throughout RL):
qwen3.6-35B-A3B-ar-sft.| checkpoint | FVE (norm) | cos mean / median | FVE (raw) |
|---|---|---|---|
| SFT (av-sft) | 0.418 | 0.807 / 0.835 | −0.26 |
| GRPO 200 | 0.583 | 0.862 / 0.887 | +0.17 |
| GRPO 1200 (this) | 0.644 | 0.882 / 0.907 | +0.32 |
nla_meta.yaml + injection_scale.txt
included; scan for the marker token id at inference time (never precompute
positions). See the NLA repo for injection utilities.