Pre-encoded training tensors for the AnyCam camera-controlled multi-view pipeline. 53,556 multishot
samples; each is one data.pth with four VAE latents, shape (1,16,21,60,104) bf16:
video_cond_latents
source view video (multishot: shot-A[:40] + shot-B[40:], pixel-stitched)
pose_cond_latents
target canonical-space normal (geometry)… See the full description on the dataset page:
https://huggingface.co/datasets/ykeee/anycam-syncam-canon-latents.