Views
No views yet
FP8_DYNAMIC scheme in the vLLM compressed-tensors format: linear weights are stored as FP8 with per-channel scales and input activations are quantized dynamically per token. The vision tower, lm_head, and token embeddings remain in their original precision. Quantization was data-free and did not modify the merged source checkpoint.