Vision-OPD-6K: Training Data for Vision-OPD
Overview
Vision-OPD proposes a regional-to-global self-distillation framework that transfers the model's own privileged regional perception to its full-image policy, without external teacher models, ground-truth labels, reward verifiers, or inference-time tool use.
Vision-OPD instantiates two conditional policies from the same MLLM: