For 3D visual grounding in monocular RGB images using language descriptions with both appearance and geometry information,
the Mono3DRefer contains 3D object targets with their corresponding geometric text descriptions, generated by ChatGPT and refined manually.
Specifically, this dataset samples 2025 frames of images from the original KITTI, containing 41,140 expressions in total and a vocabulary of 5,271 words.
The original data can be downloaded from Google… See the full description on the dataset page:
https://huggingface.co/datasets/JadeRay-42/Mono3DRefer-back.