This quant was produced with an experimental quantization workflow intended to improve the
quality/size tradeoff relative to a conventional quant of the same source model.
Model
Base model:Blackfrost-AI/Muse-Glimmer-30B-Abliterated-BF16
Format: GGUF
Approximate file size: 11.9 GB
Quantization processing time: ~2.3 hours
Target runtimes: llama.cpp and compatible GGUF frontends
License: Apache-2.0, following the upstream model
Recommended vram A 16 Gigabyte card should run this comfortably.
Local benchmark results
The following results are from local factual-summary testing and are not a standardized
leaderboard evaluation.
Variant
Approx. size
Local factual score
Q8 reference
—
~72%
Standard IQ4_XS
~14.2 GB
~67%
This custom quant
~11.9 GB
~72%
In this local test, the custom quant matched the approximate Q8 factual score while being
about 2.3 GB smaller than the tested standard IQ4_XS representation.
Results may vary with prompt, sampling settings, runtime, hardware, and benchmark methodology.