Per-step margin summary statistics exported from a margin-DPO training run.
Model repo id: W-61/llama-3-8b-base-margin-dpo-4xh100-real
Base model: princeton-nlp/Llama-3-Base-8B-SFT
Run name: llama-3-8b-base-margin-dpo-4xh100-real
Margin log path: outputs/llama-3-8b-base-margin-dpo-4xh100-real/margin_logs
Published split: train
Rows: 48
epoch
step
batch_size
mean
std
min
p10
median
p90
max… See the full description on the dataset page:
https://huggingface.co/datasets/W-61/llama-3-8b-base-margin-dpo-4xh100-real-margin.