Each example corresponds to one training run, including the training configuration and performance metrics (C4-en evaluation loss for Marin, and smoothed pretraining loss for StepLaw).
marin = load_dataset("zhqwqwq/NCPL-Pretraining-Logs", "marin", split="train")
steplaw =… See the full description on the dataset page:
https://huggingface.co/datasets/zhqwqwq/NCPL-Pretraining-Logs.