Views
No views yet
mm_projector.bin, processor / tokenizer, effective_config.json
(no QK-norm sidecar).checkpoint-4361/ — the final Stage 1 HF Trainer checkpoint:
model-*.safetensors (~7.7 GB, sharded) — full LM + projector at step 4361.global_step4361/ — DeepSpeed ZeRO-2 optimizer + bf16 master weights.trainer_state.json, training_args.bin, rng_state_*.pth.zero_to_fp32.py — DeepSpeed helper to consolidate the ZeRO shards.checkpoint-4361/:1torchrun --nproc_per_node=N -m training.train.cli \
2 --config training/configs/3b/c3_vanilla_align.yaml \
3 --resumeoutput_dir set to a directory that contains checkpoint-4361/.