Views
No views yet
step125750.ckpt
base, used to sanity-check the training pipeline. It exhibits a confirmed collapse failure mode
(action-image head produces near-blank/constant output) and should not be used for inference or
as a starting point for further training without addressing the root cause.anyeZHY/ActionImages official checkpoint--dataset_name droid@0.4,bridge@0.3,rlbench@0.3 (mixture, but each step cycles through
ONE fixed sample per dataset — not real random mixing)--action_interface pix --full_param False (2.57B trainable params: cam_encoder, projector,
self_attn, text_embedding)--max_steps 1000 --learning_rate 1e-5 --lr_scheduler_type constant--per_device_train_batch_size 1, 3x GPU (DeepSpeed ZeRO-2 + CPU offload)r_peak (heatmap peak intensity, 0-255 scale) is 2-5 across
all evaluated episodes and both camera views, vs. 150-230 for the official checkpoint on the
same episodes. a2v-conditioned video generation is also degraded (PSNR 2-7dB below official).ttd/plan/data_pipeline_and_migration.md §8) for the full input/output/loss breakdown.final_model.ckpt — 12.8GB, DeepSpeed-merged full state dict after 1000 steps.