VAMOS Navigation Dataset (with visual-language co-training data) combines multiple publicly available datasets and in-domain Spot data collected by the authors.This version includes navigation data as well as co-training data from COCO-QA and Localized Narratives.
The dataset includes:
100% of TartanDrive 2 data
50% of SCAND data
25% of CODa data
100% of in-domain Spot data (collected by the authors)
COCO-QA and Localized Narratives… See the full description on the dataset page: https://huggingface.co/datasets/mateoguaman/vamos_dataset.