This model is a fine-tuned version of
deepseek-ai/DeepSeek-R1-Distill-Llama-8B, loaded via Unsloth in 4-bit as
unsloth/DeepSeek-R1-Distill-Llama-8B-unsloth-bnb-4bit. It has been trained on customer service and general chat datasets:
The training was performed in three steps, and the final weights were merged with the base model and pushed here.
This model is released under the MIT license, allowing free use, modification, and further fine-tuning.
All code and instructions for further fine-tuning, inference, and pushing to the Hugging Face Hub are available in the open-source GitHub repository:
https://github.com/Aeshp/deepseekR1finetune
For all usage instructions, fine-tuning guides, and code, please see the
GitHub repository.