To support the visual impaired person, there are several tools including the cane.
I belive the LLM with the vision model can help.
Dataset
The road image dataset is from AI Hub.
Among the data, images in Bbox_3_new.zip have been annotated.
Description and Conversations
Based on the locations of the obstables in the image, the description has been generated.
Based on the description, the multi-turn conversation has been generated.