This model was trained on a special ChatML format with 8k context.
There are 3 different response tokens that can be used.
You can use this one if you want a medium length response. ( Greater than 64 and less than 256 tokens )
This one if you want a short length response. ( Less than 64 tokens )
This one if you want a long response. ( Greater than 256 tokens )
This llama model was trained 2x faster with
Unsloth and Huggingface's TRL library.