Views
No views yet
blk.N.attn_output.weight, where N is 10 to 35 inclusive) quantized to Q4_0, while "UDmerge-Q4_K_XXL" quantizes them to Q8_0. The latter improves refusal rate by a lot, while basically not affecting TG speed (your mileage may vary).| - | QAT BF16 | QAT Q4_0 | QAT Q4_K_M | QAT UDmerge-Q4_K_XL | QAT UDmerge-Q4_K_XXL | - | PTQ BF16 | PTQ Q4_0* | PTQ Q4_K_M |
|---|---|---|---|---|---|---|---|---|---|
| Size (GB) | 62.5** | 17.7 | 18.7 | 17.3 | 18.0 | - | 62.5** | 17.7 | 18.7 |
| PPL | 2.752 | 2.631 | 2.648 | 2.690 | 2.767 | - | 3.236 | 3.432 | 3.367 |
| KLD | 0.0000 | 0.0170 | 0.0176 | 0.0056 | 0.0017 | - | 0.0000 | 0.0714 | 0.0750 |
| Refusal | 20% | 32% | 32% | 34% | 20% | - | 16% | 16% | 15% |
| MMLU-val | 84.13% | 85.37% | 85.37% | 85.83% | 84.19% | - | 86.02% | 85.50% | 85.30% |
| MMLU-val %flips | 0.00% | 2.94% | 2.55% | 2.61% | 0.20% | - | 0.00% | 3.00% | 3.33% |
| MMLU-val %allflips | 0.00% | 3.59% | 3.00% | 3.07% | 0.26% | - | 0.00% | 3.72% | 4.05% |
*: Quant made with importance matrix ("imatrix"), results may be unreliable**: Includes multimodal ("mmproj") weights, which is ~1.2GB in BF16mlabonne/harmless_alpaca dataset's test split. Note that the dataset is processed differently, thus the numbers here are only meaningful for comparsions in this table, not with other models.mlabonne/harmful_behaviors dataset's test split. The test script, however, is adapted from Heretic to support testing needs. Note that the original author claimed 11% refusal rate for 31B and 26B-A4B models, and 6%~7% for 12B, which is not reproduced here; this is probably due to test method differences, but please take the numbers here with a grain of salt.cais/mmlu dataset's validation split (1531 questions). All tests are done once with temperature 0.0 and reasoning off. MTP is not enabled during testing. See the test script and raw data for details.| - | QAT BF16 | QAT Q4_0 | QAT Q4_K_M | QAT UDmerge-Q4_K_XL | QAT UDmerge-Q4_K_XXL | - | PTQ BF16 | PTQ Q4_0* | PTQ Q4_K_M |
|---|---|---|---|---|---|---|---|---|---|
| MMLU-val-v1 | 84.78% | 85.63% | 85.56% | 85.89% | 85.11% | - | 86.35% | 85.30% | 85.89% |
| MMLU-val-v1 %nulls | 0.20% | 0.72% | 0.46% | 0.52% | 0.20% | - | 0.46% | 0.52% | 0.46% |
| MMLU-val-v1 %flips | 0.00% | 2.29% | 1.44% | 1.76% | 0.46% | - | 0.00% | 3.27% | 2.94% |
| MMLU-val-v1 %allflips | 0.00% | 3.14% | 2.16% | 2.42% | 0.59% | - | 0.00% | 4.11% | 3.40% |