Bit-packed to INT4 from the raw OSTQuant W4A4KV4 checkpoint. Cannot be loaded
via vanilla AutoModelForCausalLM.from_pretrained(...) — the OSTQuant runtime
is required to unpack and evaluate correctly.
Compressed from ~8.9 GB (raw .bin) to ~2.9 GB (3× compression).
Weights on 4-bit grid. Activations AND KV cache also quantized to 4 bits at inference.
Reference PPL when loaded via OSTQuant runtime: 17.02 (WikiText-2 seq 2048, raw .bin).