This model is a fine-tuned version of
meta-llama/Llama-3.2-1B-Instruct, fine-tuned using DeepSpeed's ZeRO-2 Offload to enable efficient training of billion-parameter models on single GPU setups.
This model was fine-tuned using the
arxiv-abstract-dataset on a single Tesla P100 16GB GPU with CPU memory offloading.
For detailed implementation, DeepSpeed configuration, and hardware requirements, please check out the
project repository.