4-bit (Q4_K_M) GGUF quantization of wrtdevcod/fintune-qwen2.5-1.5b-lora,
a LoRA fine-tune of Qwen2.5-1.5B-Instruct for financial sentiment classification.
Original size (f16): 2.9 GB
Quantized size (Q4_K_M): 941 MB
Benchmark: 88.9% accuracy on 486-example held-out test set (base model: 50.6%)
CPU-friendly, works with llama.cpp / llama-cpp-python.