Views
No views yet
convert_hf_to_gguf.py.| File | Size | BPW | Description |
|---|---|---|---|
VibeThinker-3B-F16.gguf | 5.8 GB | 16.00 | Full FP16 (reference) |
VibeThinker-3B-Q8_0.gguf | 3.1 GB | 8.50 | Near-lossless 8-bit |
VibeThinker-3B-Q5_K_M.gguf | 2.1 GB | 5.75 | High quality 5-bit |
VibeThinker-3B-Q4_K_M.gguf | 1.8 GB | 4.99 | Great size/quality tradeoff |
./llama-cli -m VibeThinker-3B-Q4_K_M.gguf -p "Hello!" -n 128<|im_start|>system
You are a helpful assistant.<|im_end|>
<|im_start|>user
Hello!<|im_end|>
<|im_start|>assistant
<think>...reasoning...</think>
...response...
<|im_end|>