<|begin_of_text|><|start_header_id|>system<|end_header_id|> You are a helpful, respectful and honest assistant. Always answer as helpfully as possible. If a question does not make any sense, or is not factually coherent, explain why instead of answering something not correct. If you don't know the answer to a question, please don't share false information. <|eot_id|>
big thanks to @mlabonne for abliterating llama 3.1 8b, and creating this wonderful model: mlabonne/Meta-Llama-3.1-8B-Instruct-abliterated
all future versions of MFANN will use mlabonne's model as a base!
Note: You can also use this checkpoint directly through the usage steps listed in the Llama.cpp repo as well.
Step 1: Clone llama.cpp from GitHub.
git clone https://github.com/ggerganov/llama.cpp
Step 2: Move into the llama.cpp folder and build it with LLAMA_CURL=1 flag along with other hardware-specific flags (for ex: LLAMA_CUDA=1 for Nvidia GPUs on Linux).
cd llama.cpp && LLAMA_CURL=1 make
Step 3: Run inference through the main binary.
./llama-cli --hf-repo netcat420/MFANNv0.19-Q4_K_M-GGUF --hf-file mfannv0.19-q4_k_m.gguf -p "The meaning to life and the universe is"