We introduce IDKVQA, an embodied dataset specifically designed and annotated for visual question answering using the agent’s observations during navigation,
where the answer includes not only Yes and No, but also I don’t know.
Please see our ICCV 25 accepted paper: Collaborative Instance Object Navigation: Leveraging Uncertainty-Awareness to Minimize Human-Agent Dialogues
For more… See the full description on the dataset page:
https://huggingface.co/datasets/ftaioli/IDKVQA.