A large-scale community-contributed robotics dataset for vision-language-action learning, featuring 128 datasets from 55 contributors worldwide.
We used this dataset to pretrain SmolVLA. However, this is not a complete set, but the dataset that we selected using specific filters, like fps, min num of episodes, and some qualitative assessment of video qualities, using the
https://huggingface.co/spaces/Beegbrain/FilterLeRobotData tool. We also manually curated the… See the full description on the dataset page:
https://huggingface.co/datasets/HuggingFaceVLA/community_dataset_v1.