Some GGUF v2 quantizations of the model
princeton-nlp/Sheared-LLaMA-1.3B
Sheared-LLaMA-1.3B is a model pruned and further pre-trained from
meta-llama/Llama-2-7b-hf. We dynamically load data from the
RedPajama dataset. We use 0.4B tokens for pruning and 50B tokens for continued pre-training the pruned model.
We evaluate on an extensive set of downstream tasks including reasoning, reading comprehension, language modeling and knowledge intensive tasks. Our Sheared-LLaMA models outperform existing large language models.