EgoCOT is a large-scale embodied planning dataset, which selected egocentric videos from the Ego4D dataset and corresponding high-quality step-by-step language instructions, which are machine generated, then semantics-based filtered, and finally human-verified.
For mored details, please visit EgoCOT_Dataset.
If you find this dataset useful, please consider citing the paper,
@article{mu2024embodiedgpt,
title={Embodiedgpt: Vision-language pre-training via embodied chain of thought}… See the full description on the dataset page:
https://huggingface.co/datasets/wofmanaf/ego4d-video.