Finetuned from model : unsloth/phi-3.5-mini-instruct-bnb-4bit
** Trained on the Open SFT o1 Ultra Data set , quality 10 Q&A pairs only. This has given it the ability for some CoT style reasoning. your mileage may vary.
This is still WiP.
This llama model was trained 2x faster with Unsloth and Huggingface's TRL library.