This repository contains the final merged
SFT+GRPO DQ-Pilot checkpoint from
DialogueVPR: Towards Conversational Visual Place Recognition.
DQ-Pilot examines candidate street-view images and the dialogue history, then asks a
discriminative follow-up question to improve place retrieval.
1from transformers import AutoProcessor, Qwen2_5_VLForConditionalGeneration
2
3model_id = "graysonggg/dlgpr"
4processor = AutoProcessor.from_pretrained(model_id)
5model = Qwen2_5_VLForConditionalGeneration.from_pretrained(
6 model_id, torch_dtype="auto", device_map="auto"
7)
The complete CMPL retrieval and dialogue pipeline is available in the
official code repository. Training artifacts are
available in the associated
Hugging Face data repository.
1@inproceedings{song2026dialoguevpr,
2 title={DialogueVPR: Towards Conversational Visual Place Recognition},
3 author={Song, Yukun and Wang, Changwei and Pei, Xingtian and Xu, Shibiao and Xu, Wenhao and Chen, Shunpeng and Zhang, Yu and Zhang, Ke and Xu, Rongtao and Feng, Xuxiang and Wang, Pengyang},
4 booktitle={Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition},
5 pages={41100--41110},
6 year={2026}
7}
8
9@misc{song2026dialoguevpr_arxiv,
10 title={DialogueVPR: Towards Conversational Visual Place Recognition},
11 author={Song, Yukun and Wang, Changwei and Pei, Xingtian and Xu, Shibiao and Xu, Wenhao and Chen, Shunpeng and Zhang, Yu and Zhang, Ke and Xu, Rongtao and Feng, Xuxiang and Wang, Pengyang},
12 year={2026},
13 eprint={2607.14115},
14 archivePrefix={arXiv},
15 primaryClass={cs.AI},
16 url={https://arxiv.org/abs/2607.14115}
17}