Preprocessed (FastVideo parquet) data used to train the TrackWan point-track-conditioned
video model. Derived from OpenVid; each row is one 121-frame clip with everything the
trainer consumes, so it is self-contained (no re-encoding needed).
vae_latent — Wan VAE latent of the clip
first_frame_latent — I2V conditioning… See the full description on the dataset page:
https://huggingface.co/datasets/noctuashap/openvid-wantrack-processed.