Qwen3.8-27B-AEON-Ultimate-Uncensored — GGUF (UD Quants)
Unsloth Dynamic-style (UD) GGUF quantizations of
AEON-7/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-BF16
Two quant families are available:
- AWQ-UD (recommended): AWQ channel pre-scaling applied before quantization. Lower perplexity than baseline at every bit width. All AWQ-UD quants beat F16 on perplexity.
- Baseline UD: Standard quantization without pre-scaling.
Every quant uses per-tensor overrides (sensitivity-driven) + importance matrix (multi-domain calibration). All SSM recurrence tensors are preserved at source precision. MTP speculative decoding and vision (mmproj) are preserved.
Quant Comparison
AWQ-UD (recommended)
AWQ pre-scaling (256 samples x 512 tokens, W4A16_ASYM target) redistributes weight magnitudes before quantization. All AWQ-UD quants beat F16 reference perplexity.
| File | Quant | Size | PPL | KL mean | tg t/s |
|---|
| F16 (reference) | F16 | 50.9 GB | 2.8950 | — | 30.3 |
| AWQ-UD-Q8_0 | Q8_0 | 34.7 GB | 2.8807 | 0.00384 | 42.4 |
| AWQ-UD-Q6_K | Q6_K | 30.6 GB | 2.8795 | 0.00376 | 47.2 |
| AWQ-UD-Q5_K_M | Q5_K_M | 28.7 GB | 2.8672 | 0.00945 | 49.0 |
| AWQ-UD-IQ4_XS | IQ4_XS | 25.9 GB | 2.8887 | 0.01922 | 53.2 |
Baseline UD
| File | Quant | Size | PPL | KL mean | tg t/s |
|---|
| UD-Q8_0 | Q8_0 | 34.7 GB | 2.8918 | 0.00181 | 42.4 |
| UD-Q6_K | Q6_K | 30.6 GB | 2.8891 | 0.00153 | 47.2 |
| UD-Q5_K_M | Q5_K_M | 28.7 GB | 2.8856 | 0.00822 | 49.0 |
| UD-IQ4_XS | IQ4_XS | 25.9 GB | 2.8990 | 0.01791 | 53.2 |
Benchmarked on
NVIDIA RTX PRO 6000 Blackwell (96 GB VRAM), llama.cpp (
a4501150/llama.cpp), pp=512, tg=128. Throughput is identical between AWQ-UD and baseline UD at the same bit width because AWQ does not change tensor sizes.
AWQ-UD vs baseline: AWQ wins on perplexity (0.010-0.018 lower PPL). Baseline wins on KL divergence (closer to F16 output distribution). AWQ changes the channel basis, which shifts the output distribution away from F16, but absolute quality improves.
SGLang Throughput (NVFP4 + DFlash2)
For maximum throughput, serve the NVFP4 checkpoint via SGLang with DFlash2 speculative decoding:
| Config | Single-user t/s | 3 users (agg) | 6 users (agg) |
|---|
| llama.cpp NVFP4 | 64 | — | — |
| llama.cpp UD-Q6_K + DSpark | 72 | — | 178 (5 users) |
| SGLang NVFP4 + DFlash2 | 146-178 | 391 | 493 |
SGLang is 2.3-2.8x faster single-user and 2.8x faster concurrent vs llama.cpp.
What Makes These Different
AWQ Pre-Scaling (AWQ-UD only)
AWQ (Activation-Aware Weight Quantization) applies per-channel scaling to redistribute weight magnitudes before quantization. This makes outlier channels less damaging when quantized. The scaling is applied at BF16 precision and is lossless — the model produces identical output before quantization. After scaling, the full GGUF pipeline runs: convert, importance matrix, sensitivity analysis, quantize with per-tensor overrides.
SSM Recurrence Preservation
Qwen3.8 is a hybrid GatedDeltaNet + attention model. 48 of 64 layers use a recurrent SSM where quantization error compounds across token positions. All SSM recurrence tensors are preserved at source precision (F16) — never quantized.
| Tensor | Count | Precision | Rationale |
|---|
ssm_alpha, ssm_beta | 96 | F16 | State update projections — error accumulates in recurrence |
ssm_out | 48 | F16 | Output projection feeds directly into residual stream |
ssm_a, ssm_conv1d, ssm_dt, ssm_norm | 192 | F32 | Small state tensors (llama-quantize keeps 1D/small tensors at F32) |
attn_qkv (SSM input projection) | 48 | F16 | Highest measured KL sensitivity |
attn_gate (SSM gate projection) | 48 | F16 | Second-highest measured KL sensitivity |
Per-Tensor Sensitivity Analysis
Each tensor group was probed by quantizing only that group to Q4_0 while keeping the rest at F16, then measuring KL divergence. The override generator assigns precision based on measured sensitivity:
| Precision | Tensor Groups | Override Count |
|---|
| F16 | SSM recurrence, norms, biases, MTP layer | 512 |
| F16 | All attention tensors (attn_qkv, attn_gate, attn_v, attn_q, attn_k, attn_output), ffn_down edge | 173 |
| Base quant | FFN middle layers, FFN edge gate/up, embeddings | ~181 |
685 total overrides — every non-FFN tensor has an explicit precision assignment. No dependence on llama-quantize's internal promotion rules.
Multi-Domain Calibration + GPU Imatrix
Calibrated on a balanced mix across 4 domains from 13 HF datasets:
| Domain | Token Budget | Sources |
|---|
| General | 1M | ultrachat, OpenHermes, COIG-CQIA, LongAlpaca, pg19, froggeric/imatrix |
| Code | 750K | Magicoder-Evol-Instruct-110K |
| Reasoning | 750K | OpenMathInstruct-2, OpenR1-Math-220k |
| Agentic | 500K | glaive-function-calling-v2, xlam-function-calling-60k, hermes-function-calling-v1 |
Special tokens from source datasets are stripped automatically. Samples are kept whole — never truncated mid-conversation.
The importance matrix is generated with a PyTorch GPU-native generator at 32,768 context — uses forward hooks to accumulate squared activations on GPU with zero PCIe D2H copies. Supports multi-GPU via device_map="auto".
Per-domain imatrices are merged with equal weights (DI-MATRIX approach).
MTP + Vision Preserved
- MTP (Multi-Token Prediction): Draft head (blk.64) pinned at F16. Use
--spec-type draft-mtp --spec-draft-n-max 3 for ~1.5-2x faster generation.
- Vision: mmproj file contains the full vision encoder. Use
--mmproj flag with llama-server for image/video understanding.
Files
| File | Description | Size |
|---|
Qwen3.8-27B-AEON-AWQ-UD-Q8_0.gguf | AWQ — Highest quality | 34.7 GB |
Qwen3.8-27B-AEON-AWQ-UD-Q6_K.gguf | AWQ — Recommended — best quality/size | 30.6 GB |
Qwen3.8-27B-AEON-AWQ-UD-Q5_K_M.gguf | AWQ — Balanced | 28.7 GB |
Qwen3.8-27B-AEON-AWQ-UD-IQ4_XS.gguf | AWQ — Smallest | 25.9 GB |
Qwen3.8-27B-AEON-UD-Q8_0.gguf | Baseline — Highest quality | 34.7 GB |
Qwen3.8-27B-AEON-UD-Q6_K.gguf | Baseline — Best quality/size | 30.6 GB |
Qwen3.8-27B-AEON-UD-Q5_K_M.gguf | Baseline — Balanced | 28.7 GB |
Qwen3.8-27B-AEON-UD-IQ4_XS.gguf | Baseline — Smallest | 25.9 GB |
Qwen3.8-27B-AEON-mmproj-F16.gguf | Vision encoder (use with --mmproj) | 885 MB |
Qwen3.8-27B-sharp.jinja | Enhanced chat template (terse output, reasoning effort, tool error detection) | 18 KB |
imatrix_merged.dat | Importance matrix for requantization | 13 MB |
Usage
llama-server (recommended)
1# AWQ-UD-Q6_K with sharp template, MTP + vision
2llama-server \
3 -m Qwen3.8-27B-AEON-AWQ-UD-Q6_K.gguf \
4 --mmproj Qwen3.8-27B-AEON-mmproj-F16.gguf \
5 -ngl 99 \
6 --flash-attn \
7 -c 262144 \
8 --parallel 3 \
9 -kvu \
10 --jinja \
11 --chat-template-file Qwen3.8-27B-sharp.jinja \
12 --reasoning-format deepseek \
13 --reasoning-preserve \
14 --spec-type draft-mtp \
15 --spec-draft-n-max 3 \
16 --host 0.0.0.0 --port 8080
Note: --spec-type draft-mtp requires llama.cpp b9375+.
--reasoning-format deepseek extracts thinking into
message.reasoning_content in API responses. The sharp template (by
froggeric) enables thinking by default with terse output, reasoning effort control, and tool call error detection.
llama-cli
1llama-cli \
2 -m Qwen3.8-27B-AEON-AWQ-UD-Q6_K.gguf \
3 -ngl 99 \
4 --flash-attn \
5 -c 262144 \
6 --jinja \
7 --chat-template-file Qwen3.8-27B-sharp.jinja \
8 --reasoning on \
9 --reasoning-preserve
Sampling Parameters
| Mode | temperature | top_p | top_k | min_p | presence_penalty | repetition_penalty |
|---|
| Thinking (default) | 1.0 | 0.95 | 20 | 0.0 | 0.0 | 1.0 |
| Non-thinking | 0.7 | 0.80 | 20 | 0.0 | 1.5 | 1.0 |
Do not use greedy decoding (temperature=0). reasoning_effort controls thinking depth independently: xhigh (default), medium, low.
Architecture
Qwen3.8-27B is a hybrid SSM-attention model:
- 64 transformer layers + 1 MTP layer (blk.0-64)
- 48 SSM layers (GatedDeltaNet, no KV cache) + 16 full attention layers (every 4th layer)
- 27B parameters, 24 attention heads, 4 KV heads, head dim 256
- Vocab: 248,320 tokens, native context: 262,144 tokens
Quantization Pipeline
Built with super-quant:
- AWQ pre-scaling (AWQ-UD only): Apply per-channel weight scaling via llm-compressor AWQModifier (256 calibration samples, W4A16_ASYM target). Strip all quantization state. Save as plain BF16.
- Convert HF to F16 GGUF (with MTP tensors) + mmproj GGUF (vision)
- Multi-domain calibration data from 13 HF datasets, special tokens stripped
- GPU-native importance matrix generation (PyTorch, 32k context) + weighted merge
- Per-tensor sensitivity analysis (KL divergence probing against F16 logits)
- Hybrid override generation — SSM at source precision, sensitivity-driven for the rest
- Quantize with per-tensor overrides + imatrix
- Benchmark: throughput + perplexity + KL divergence vs F16
Links
- Base model: AEON-7/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-BF16
- Quantization pipeline: super-quant
- llama.cpp fork: a4501150/llama.cpp (DFlash, MTP fixes, Blackwell FA4)
Credits
License: Apache-2.0