Views
No views yet
rfp458-pack-quantized): iq4_nl non-uniform 4-bit codebook, group size 16, signed-int8 block mantissa + per-channel int8 exponent, with hadamard16 weight rotation.in_proj_qkv / in_proj_z / out_proj, and the full self-attention q/k/v/o projections. Embeddings, lm_head, the GDN gating projections (in_proj_a / in_proj_b), conv1d, norms, and the entire vision tower are kept in bf16.| Build | Size | PPL |
|---|---|---|
| RFP458 (this model) | 20.5 GB | 6.936 |
| FP8 (RedHatAI/Qwen3.6-27B-FP8) | ~27 GB | 7.071 |