PDS-DPO-7B is a vision-language model built upon LLaVA 1.5 7B and trained using the proposed Preference Data Synthetic Direct Preference Optimization (PDS-DPO) framework. This approach leverages synthetic data generated using generative and reward models as proxies for human preferences to improve alignment, reduce hallucinations, and enhance reasoning capabilities.
1@article{wijaya2024multimodal,
2 title={Multimodal Preference Data Synthetic Alignment with Reward Model},
3 author={Wijaya, Robert and Nguyen, Ngoc-Bao and Cheung, Ngai-Man},
4 journal={arXiv preprint arXiv:2412.17417},
5 year={2024}
6}