I think this is the original one I uploaded when trying to do continual pre-training but couldn't get the LoRA adaptor to merge with the base model so I just uploaded it to merge later. If I remember correctly, it's bad.
This gemma3 model was trained 2x faster with
Unsloth and Huggingface's TRL library.