Views
No views yet
ReflectiVA), utilizes reflective tokens to dynamically determine the need for external knowledge
and predict the relevance of information retrieved from an external database.
Tokens are trained following a two-stage two-model training recipe. This ultimately enables the MLLM to manage external knowledge
while preserving fluency and performance on tasks where external knowledge is not needed.ReflectiVA for knowledge-based visual question answering, highlighting its
superior performance compared to existing methods.ReflectiVA.1@inproceedings{cocchi2024augmenting,
2 title={{Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering}},
3 author={Cocchi, Federico and Moratelli, Nicholas and Cornia, Marcella and Baraldi, Lorenzo and Cucchiara, Rita},
4 booktitle={Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition},
5 year={2025}
6}