Per-step margin summary statistics exported from a New-DPO training run.
Model repo id: jackf857/llama-3-8b-base-new-dpo-ultrafeedback-4xh200-batch-128-q_t-0.5-s_star-0.4-20260429-032138
Base model: /scratch/qu.yang1/dynamic-dpo-v4/base_models/llama-3-8b-base-sft-ultrachat-8xh200
Training run name:… See the full description on the dataset page:
https://huggingface.co/datasets/jackf857/llama-3-8b-base-new-dpo-ultrafeedback-4xh200-batch-128-q_t-0.5-s_star-0.4-20260429-032138-margin.