This dataset documents blind spots for the pretrained video world model facebook/vjepa2-vith-fpc64-256.
The model produces nearly identical embeddings for videos whose temporal order has been severely corrupted, indicating weak sensitivity to temporal directionality and causal motion structure.
Model name: facebook/vjepa2-vith-fpc64-256Type: Self-supervised video world model (JEPA-style joint embedding predictive… See the full description on the dataset page:
https://huggingface.co/datasets/Nuntea/vjepa2-temporal-order-blindspots.