A Multimodal Dataset for Vision-Language Grounding in Human-Robot Interaction (HRI)๐ Language: Japanese | ๐ค Focus: Embodied AI | ๐ Size: 142 Scenes | ๐ฌ Granularity: Object-Level
J-ORA (Japanese Object Reference and Action) is a multimodal benchmark for grounded vision-language learning in robotics. It is designed for understanding Japaneseโฆ See the full description on the dataset page:
https://huggingface.co/datasets/atamiles/J-ORA.