Views
No views yet
| Quant | Recommended usage |
|---|---|
| Q2_K | Very low memory devices |
| Q3_K_M | Mobile devices |
| Q4_K_M | Balanced quality and speed |
| Q5_K_M | Higher quality |
| Q6_K | Strong quality-performance balance |
| Q8_0 | Maximum quality |
1from llama_cpp import Llama
2
3llm = Llama.from_pretrained(
4 repo_id="Tesleum/Shirdel-Finance-E4B",
5 filename="Shirdel-Finance-E4B.Q4_K_M.gguf"
6)
7
8response = llm.create_chat_completion(
9 messages=[
10 {
11 "role":"user",
12 "content":"Explain Bitcoin halving."
13 }
14 ]
15)
16
17print(response)curl -LsSf https://llama.cpp/install.sh | shllama-cli -hf Tesleum/Shirdel-Finance-E4B:Q4_K_Mllama-server -hf Tesleum/Shirdel-Finance-E4B:Q4_K_Mollama run hf.co/Tesleum/Shirdel-Finance-E4B:Q4_K_M1FROM ./Shirdel-Finance-E4B.Q4_K_M.gguf
2
3TEMPLATE """{{ .Prompt }}"""
4
5PARAMETER temperature 0.7
6PARAMETER top_p 0.95
7PARAMETER num_ctx 1310721ollama create shirdel-finance -f Modelfile
2ollama run shirdel-finance