These models were pretrained with the spatiotemporal MAE algorithm on ~5k hours of curated human-like video data (mostly egocentric, temporally extended, continuous video recordings) and then, optionally, finetuned on various downstream tasks with few-shot supervised training.
Please note that this repository only stores the pretrained model checkpoints. Please use
the associated GitHub repository to actually load the models and use them (the GitHub repository contains detailed instructions for loading and using the models).