Views
No views yet
pack-quantized), with the lm_head untied from the embeddings and quantized to 8-bit to shrink
the per-step output-projection read. Built for fast single-request (batch-1, greedy) inference and
designed to pair with the DFlash speculative-decoding drafter.llm-compressor GPTQModifier (4-bit symmetric int, group 128,
actorder: static), calibrated on 840 public, train-split samples (ARC-Challenge, QASC,
MedMCQA, MedQA-USMLE, AQuA-RAT, CommonsenseQA, Dolly). The benchmark datasets (GPQA-Diamond,
MMLU-Pro, IFEval) are explicitly excluded from calibration.