Views
No views yet
π PAPER
π€ CLIMBLAB
π€ CLIMBMIX Figure 1: Continuously training a 1B model yields a 2.0% improvement over Llama-3.2-1B, demonstrating a more efficient scaling trend compared to prior models.
Figure 2: Pre-training a 1B model from scratch on ClimbMix shows better scaling effects than training on other datasets.β¦ See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-ClimbLab.