Hallucination-aligned visual negative construction dataset for multimodal DPO training.
Dataset Summary
This dataset contains 8,456 preference pairs with 7,971 generated negative images for cross-modal DPO training on VLMs. Each negative image is generated to visually depict the hallucinated content from the model's rejected response.