This repository contains GGUF format quantizations of the model
DavidAU/LFM2-8B-A1B-Instruct-Quantum-IQ1C-Qwen3.6-35B-A3B-DISTILL.
The following benchmarks were obtained running the model locally on a standard laptop configuration, without GPU acceleration.
-
LFM2-8B-A1B-Instruct-QwenDistill-Q4_K_M.gguf :
- Prompt Evaluation: ~41 t/s
- Token Generation: ~24 t/s
-
LFM2-8B-A1B-Instruct-QwenDistill-Q5_K_M.gguf :
- Prompt Evaluation: ~28 t/s
- Token Generation: ~20 t/s