Llama-HybridDiffusion-2B run 1 — resumable training state
Built with Llama.
This public repository is a disaster-recovery snapshot for a research reproduction of
HybridDiffusion using Qwen3.5-2B. It is not an inference-ready Transformers checkpoint.
checkpoints/step-3000/ is a complete PyTorch Distributed Checkpoint (DCP) containing
model parameters, optimizer state, learning-rate scheduler state, training state, and
rank-local dataloader state for all eight ranks.
Snapshot identity
- Training step: 3000
- DCP files: 9 (
.metadata plus eight .distcp rank shards)
- Exact directory size: 24,636,633,882 bytes
- Training world size: 8
- Base model:
Qwen/Qwen3.5-2B
Exact resumption additionally requires the matching HybridDiffusion source snapshot,
training configuration, processed datasets, and a compatible software environment.
Those materials are not included in this checkpoint repository.
Data provenance warning
The run uses a processed mixture derived from NVIDIA Nemotron releases and the
Llama-Nemotron post-training dataset. Processed Arrow data are not mirrored here
because the local transformed output no longer retains every source's per-sample
licence field. Source terms include CC BY 4.0, CC BY-SA 4.0, ODC-By, the NVIDIA Open
Model License, and Llama Community Licence conditions.
Licence
No single licence replaces the applicable upstream terms. See NOTICE and LICENSES/.
The Qwen base is Apache 2.0. HybridDiffusion code is PolyForm Noncommercial 1.0.0.
Llama and NVIDIA attribution is included conservatively due to training-data lineage.
Users must review all applicable terms before use or redistribution.