Recording acknowledgements: We thank Chenrui Shi, Haoran Xu, and Shile Li for their assistance in recording the LongVILBench demonstrations.
Paper · Code and evaluation
LongVILBench is a benchmark for long-horizon visual imitation learning from real-world tabletop demonstration videos. It evaluates whether a model can infer a temporally ordered, spatially grounded action plan from a human demonstration and translate that plan into an… See the full description on the dataset page:
https://huggingface.co/datasets/cq838/LongVILBench.