This is
LFM2.5-1.2B-Instruct quantized with
llm-compressor using AWQ (asymmetric 4-bit activation-aware quantization). The model is compatible with vLLM (tested: v0.13.0). Tested with an RTX 4090.
"
buy me a kofi"
Subscribe to
The Kaitchup. This helps me a lot to continue quantizing and evaluating models for free.