Views
No views yet
<bos><start_of_turn>user
{prompt}<end_of_turn>
<start_of_turn>model
<end_of_turn>
<start_of_turn>model
| Branch | Bits | lm_head bits | VRAM (4k) | VRAM (16k) | VRAM (32k) | Description |
|---|---|---|---|---|---|---|
| 8_0 | 8.0 | 8.0 | 11.9 GB | 15.9 GB | 21.3 GB | Maximum quality that ExLlamaV2 can produce, near unquantized performance. |
| 6_5 | 6.5 | 8.0 | 10.4 GB | 14.4 GB | 19.8 GB | Very similar to 8.0, good tradeoff of size vs performance, recommended. |
| 5_0 | 5.0 | 6.0 | 8.6 GB | 12.6 GB | 18.0 GB | Slightly lower quality vs 6.5, but usable on 8GB cards. |
| 4_25 | 4.25 | 6.0 | 7.9 GB | 11.9 GB | 17.3 GB | GPTQ equivalent bits per weight, slightly higher quality. |
| 3_5 | 3.5 | 6.0 | 7.1 GB | 11.1 GB | 16.9 GB | Lower quality, only use if you have to. |
git clone --single-branch --branch 6_5 https://huggingface.co/bartowski/Smegmma-Deluxe-9B-v1-exl2 Smegmma-Deluxe-9B-v1-exl2-6_5pip3 install huggingface-hub--revision parameter. For example, to download the 6.5 bpw branch:huggingface-cli download bartowski/Smegmma-Deluxe-9B-v1-exl2 --revision 6_5 --local-dir Smegmma-Deluxe-9B-v1-exl2-6_5huggingface-cli download bartowski/Smegmma-Deluxe-9B-v1-exl2 --revision 6_5 --local-dir Smegmma-Deluxe-9B-v1-exl2-6.5