Qwen2.5-VL-7B-Instruct trained for 30 Vision-OPD optimizer steps. The student
retains a replayable random 5% subset of post-merge visual tokens, while the
EMA teacher uses the full visual-token sequence.
Training uses Vision-OPD-6K with 96 prompts and 8 on-policy rollouts per step,
a learning rate of 2e-6, and an EMA teacher update rate of 0.05.
Standard Transformers/vLLM loads the full-token model. RandomDrop5 inference
requires the pruning-aware serving code in
prune-opd.