The VPP-SFT dataset is the SFT (Supervised Fine-Tuning) dataset used in the VPP-LLaVA paper. It is designed to provide high-quality visual grounding data in a compact format for efficient model training. The repo also benifits form ChatterBox (AAAI 2025) and Genixer (ECCV 2024)
Thanks for their wonderful works.
Purpose: The VPP-SFT dataset aims to enhance the visual grounding capabilities of multimodal large language models (MLLMs) by… See the full description on the dataset page:
https://huggingface.co/datasets/wayneicloud/VPP-SFT.