Some GGUF v2 quantizations of the model
princeton-nlp/Sheared-LLaMA-2.7B
Sheared-LLaMA-2.7B is a model pruned and further pre-trained from
meta-llama/Llama-2-7b-hf. We dynamically load data from different domains in the
RedPajama dataset. We use 0.4B tokens for pruning and 50B tokens for continued pre-training the pruned model.
We evaluate on an extensive set of downstream tasks including reasoning, reading comprehension, language modeling and knowledge intensive tasks. Our Sheared-LLaMA models outperform existing large language models.