The adapter was trained with four LoRA layers, rank 8, scale 16, a maximum
sequence length of 512, and 400 continuation steps from a saved step-100
checkpoint. Final test loss was 0.820 (perplexity 2.271).
This is an MLX-format text model and is not a Transformers/PyTorch checkpoint.