CoVLA-Dataset is a dataset comprising real-world driving videos spanning more than 80 hours. This dataset leverages a novel, scalable approach based on automated data processing and a caption generation pipeline to generate accurate driving trajectories paired with detailed natural language descriptions of driving environments and maneuvers. It includes 10,000 30-second video clips, paired with trajectory targets and language annotations generated from… See the full description on the dataset page:
https://huggingface.co/datasets/turing-motors/CoVLA-Dataset.