-
Pre-training on a dataset of 1.3 trillion tokens across 16 nodes, each equipped with four AMD Instinct MI250 GPUs. This foundational phase allows the model to grasp general language patterns and knowledge.
-
Supervised Fine-Tuning (SFT), where the model was refined using a diverse array of datasets, including Tulu V2 and OpenHermes-2.5, enhancing its performance in science, coding, and mathematics tasks.
-
Direct Preference Optimization (DPO), which aligns the model's outputs with human preferences by training it on the UltraFeedback dataset. This ensures that responses are not only accurate but also resonate with user expectations.