Views
No views yet
5latin2latout| training_mode | video (no action head — action DiT is not trained) |
| latent conditioning | 5latin2latout (cond latent frames -> 2 future latent frames) |
| video EMA | saved (_video_ema_model.*, 825 tensors, fp32) |
| EMA schedule | diffusers use_ema_warmup (inv_gamma=1.0, power=0.75, cap 0.9999) |
| extraction | WAM_VIDEO_EMA_TRAIN_EXTRACT=0 — EMA is tracked+saved only, training math unchanged |
| GPUs | 8 x H200 |
checkpoint-<step>/.global_step*) is intentionally excluded, so these checkpoints are
for inference / feature extraction / probing only — they cannot resume training.model-*.safetensors + index) and processor/config
files. The EMA copy of the video DiT is included in the same shards under the
_video_ema_model.* prefix.