This dataset contains outputs from a system that searches a video using an image of a car part.
We give the system a picture (for example a hood or door), and it returns the moments in the video where that same part appears.
The system matches meaning (car parts), not exact pixels.
The video is converted into frames (1 frame per second)
A trained detector finds car parts in every frame
All… See the full description on the dataset page:
https://huggingface.co/datasets/shiniagarwal/rav4-semantic-video-index.