A vision-language dataset designed for lightweight traffic scene understanding and contextual scene depiction tasks.
This dataset was generated using knowledge distillation from the Qwen2.5-VL-7B-Instruct Vision Language Model (VLM). Each image was processed using a structured prompting strategy to generate grounded and context-aware natural language descriptions of urban traffic scenes.
The objective of this dataset is to support the development of the… See the full description on the dataset page:
https://huggingface.co/datasets/Subh775/Traffic-Perception-VL.