TinyThinkingLlama
TinyThinkingLlama is a fine tuned version of unsloth/Llama-3.2-3B-Instruct-bnb-4bit trained on a thinking/reasoning dataset.
Model Details
Model Description
TinyThinkingLlama is a fine tuned version of unsloth/Llama-3.2-3B-Instruct-bnb-4bit trained on a thinking/reasoning dataset to gain reasoning abilites with the goal of training smaller models to comprehend tasks better and make less mistakes. By teaching smaller models how to reason we can increase their usefulness and overall accuraccy and efficiency with minimal hardware requirements increasing accessibility.
- Developed by: S'Bussiso Dube
- Model type: Text/chat
- Language(s) (NLP): English
- License: MIT
- Finetuned from model unsloth/Llama-3.2-3B-Instruct-bnb-4bit
Uses
This model is fine tuned for general assistance and smaller tasks that still require some level of complex reasoning. The dataset used to fine-tune TinyThinkingLlama consists of multiple domains such as science, coding, and literature. This model is in standard GGUF format so it can be used in most places and platforms such as Ollama, HuggingFace, Llama.cpp, etc.
Direct Use
As is this model can be used for standard inference.
[More Information Needed]
Risks and Limitations
TinyThinkingLlama is still prone to mistakes as all models are and smaller models such as 3B parameters are even more prone to make mistakes. The reasoning ability only helps and should not be used for critical tasks.
[More Information Needed]
Recommendations
Do not use TinyThinkingLlama on large tasks or tasks of great importance. Do not give tool access unless in a sandboxed environment.
How to Get Started with the Model
Use the code below to get started with the model.
[More Information Needed]
Training Details
Training Data
[More Information Needed]
Training Procedure
Preprocessing [optional]
[More Information Needed]
Training Hyperparameters
Training Details:
Hyperparameters
- Epochs: 1
- Batch size: 2
- Learning rate: 0.0002
- Optimizer: AdamW 8-bit
- Max steps: 0
- Context length: 4096
- Warmup steps: 5
LoRA
-
Rank: 16
-
Alpha: 16
-
Dropout: 0
-
Variant: lora
-
Hardware Type: T4 GPU
-
Hours Trained: 3 Hours
-
Cloud Provider: Google Colab
Summary
Environmental Impact
Smaller models run on mininal hardware requirements negating the neccessity of more GPUs or more powerfull ones. By improving and using smaller models less compute and electricity is used in both training and inference lessening strain on the grid and improving environmental impact.
Model Architecture and Objective
[More Information Needed]
Model Card Contact
[More Information Needed]