Views
No views yet
python convert_hf_to_gguf_update.py <hf_token>
python convert_hf_to_gguf.py models/tokenizers/gemma-2/ --outfile models/ggml-vocab-gemma-2.gguf --vocab-only
test-tokenizer-0 models/ggml-vocab-gemma-2.gguf<start_of_turn>user
ここにpromptを書きます<end_of_turn>
<start_of_turn>model
| クオンツ | VRAM |
|---|---|
| IQ4_XS | 10GB |
| Q4_K_M | 11GB |
| Q5_K_M | 11GB |
| Q6_K | 12GB |
| Q8_0 | 14GB |
| bf16 | 22GB |
-fa オプションによるFlash Attentionの使用はできません。