Training data for Vision-OPD: Learning to See Fine Details for Multimodal LLMs via On-Policy Self-Distillation.
Vision-OPD proposes a regional-to-global self-distillation framework that transfers the model's own privileged regional perception to its full-image policy, without external teacher models, ground-truth labels, reward verifiers, or inference-time tool use.
Vision-OPD instantiates two conditional policies from… See the full description on the dataset page:
https://huggingface.co/datasets/zwyang6/Vision-OPD-6K.