Supervised fine-tuning (SFT) dataset for an autonomous-vehicle vision-language
waypoint-prediction model. Contains 2,789,773 samples from 150 000
driving scenes (18 seconds per scene, sampled at 1 Hz) recorded in the
United States.
WebDataset — 100 uncompressed .tar shards,
each containing pairs of files per sample:
{key}.json
Metadata (see schema below)… See the full description on the dataset page:
https://huggingface.co/datasets/tom-jerry-123/Physical-AI-AV-US.