We are excited to introduce the Z1 family of models! These models are based on the OLMo 2 1B architecture developed by Allen Institute for AI. Beginning with the pre-training checkpoint for OLMo 2 1B, we performed continued pre-training (i.e., midtraining) on Z1 1B Hybrid using the same dataset as OLMo 2 1B (dolmino-mix-1124).
What is unusual about the Z1 models is that the continued pre-training was performed via Zettafleet’s AI Training Platform on 8 NVIDIA GPUs in a fully decentralized way, without the use of high-bandwidth near-range communication links (i.e., NVLink) between the accelerators. See our blog post for further details.
We release the following models as part of the Z1 family:
zettafleet/z1-1b-hybrid: A base model where continued pre-training was performed in a fully decentralized way on 8 NVIDIA H100 GPUs.
zettafleet/z1-1b-hybrid-rtx: A base model where continued pre-training was performed in a fully decentralized way on 8 NVIDIA RTX Pro 6000 GPUs.
AI models can be prompted by users to generate harmful and sensitive content. Such content may also be produced unintentionally, especially in cases involving bias, so we recommend that users consider the risks when applying this technology. Additionally, many statements from Z1 or any LLM are often inaccurate, so facts should be verified.