Views
No views yet
qwen35.nextn_predict_layers=1, +~0.24 GiB vs the old trunk-only files). You do not need a second GGUF for draft-mtp.qwen35). 64 language-trunk blocks plus 1 nextn/MTP block. Converted from the AEON BF16 master (no --no-mtp). Pair a separate mmproj if you need vision.<think> blocks and breaks multi-turn agents. These GGUFs bake a fixed template. Use --jinja.llama-server -m Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-Q4_K_M.gguf --jinja --reasoning-format deepseek| File | Quant | Bytes | GiB | Notes |
|---|---|---|---|---|
Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-Q2_K.gguf | Q2_K | 10864592928 | 10.12 | live; MTP baked in |
Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-Q3_K_M.gguf | Q3_K_M | 13500737568 | 12.57 | live; MTP baked in |
Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-Q4_K_M.gguf | Q4_K_M | 16810715168 | 15.66 | live; MTP baked in |
Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-Q5_K_M.gguf | Q5_K_M | 19535702048 | 18.20 | live; MTP baked in |
Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-Q6_K.gguf | Q6_K | 22431000608 | 20.89 | live; MTP baked in; largest full-GPU on 24GB Turing |
Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-Q8_0.gguf | Q8_0 | 29047085088 | 27.05 | live; MTP baked in; will not -ngl 99 on 24GB |
draft-mtp. --parallel 1 is required.1llama-server \
2 -m Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-Q4_K_M.gguf \
3 --spec-type draft-mtp --spec-draft-n-max 2 --spec-draft-p-min 0.7 \
4 --parallel 1 --jinja --reasoning-format deepseek \
5 -a qwen38-27b-aeon --host 127.0.0.1 --port 8085 \
6 -ngl 99 -fa on -b 512 -ub 512 -c 32768mtp-* sidecars are still in the repo if you want a separate -md draft. They are not required for the files above.git clone. Grab only the file that exists.1export HF_XET_HIGH_PERFORMANCE=1
2hf download vcruz305/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-GGUF \
3 --local-dir Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-GGUF \
4 --include "Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-Q4_K_M.gguf"convert_hf_to_gguf.py --outtype f16 (MTP mixin on, no --no-mtp) → llama-quantize