Views
No views yet
track_encoder + patch-embed track slot), head
trainable, on a small 50-clip 720p synthetic set. Random track-ID sampling
(WANTRACK_FIXED_SAMPLE=0) so the pathway generalizes over arbitrary ID/point assignments
(the invariance stage — it can only be learned while the head is trainable, before stage-1
freezes it).checkpoint-2000): fixed-ID overfit → random-ID overfit,
continued 2000→3000.WANTRACK_PMASK=0 (mid-frame masking belongs to
stage-2) and WANTRACK_IMAGE_COND=0 (image cross-attention is enabled from stage-1 on the
pristine merged base; turning it on mid-overfit shocks a CLIP-free-trained seed).lr 1e-4 (constant), flow_shift 6, sparse conditioning (WANTRACK_SPARSE=1,
EXTRA_RANDOM=20), d64 track-ID embedding + bias track encoder.dcp/__<rank>_0.distcp (8 ranks, HSDP 2×4). Not a from_pretrained model — consolidate
via the FastVideo export step (03_export.sh) to get a diffusers transformer.data_pipeline/720_stage_1/ recipe.