| Nome | Método Quant | Bits | Tamanho | Desc |
|---|---|---|---|---|
| llama-2-7b-langchain-chat-q4_0.gguf | q4_0 | 4 | 3.56 GB | Quantização em 4-bit. |
| llama-2-7b-langchain-chat-q4_1.gguf | q4_1 | 4 | 3.95 GB | Quantização em 4-bit. Acurácia maior que q4_0 mas não tão boa quanto q5_0. Inferência mais rápida que os modelos q5. |
| llama-2-7b-langchain-chat-q5_0.gguf | q5_0 | 5 | 4.33 GB | Quantização em 5-bit. Melhor acurácia, maior uso de recursos, inferência mais lenta. |
| llama-2-7b-langchain-chat-q5_1.gguf | q5_1 | 5 | 4.72 GB | Quantização em 5-bit. Ainda Melhor acurácia, maior uso de recursos, inferência mais lenta. |
| llama-2-7b-langchain-chat-q8_0.gguf | q8_0 | 8 | 6.67 GB | Quantização em 8-bit. Quase indistinguível do float16. Usa muitos recursos e é mais lento. |
llama.cpp./main -m ./models/llama-2-7b-langchain-chat/llama-2-7b-langchain-chat-q5_1.gguf --color --temp 0.5 -n 256 -p "<s>[INST] Há muito tempo atrás, numa galáxia distante [/INST] Assistant Message </s>"<s>[INST] Prompter Message [/INST] Assistant Message </s>