Fine tune of "LFM2-8B-A1B" using Unsloth using custom dataset(s), 128k context in 16 bit precision.
This model is a sparse mixture of experts model (32) with 4 experts activated.
Speed exceeds 50-100 t/s on CPU // 200 t/s on most cards // 400 t/s + on 5090 at QUANT Q6K [4 experts].
One example generation below.
Can also be used on phones // mobile devices.
IN HOUSE BENCHMARKS [by Nightmedia]:
arc-c arc/e boolq hswag obkqa piqa wino
LFM2-8B-A1B-GLM-4.7-Flash-Thinking-Quantum-IQ1C
mxfp8 0.495,0.709,0.759,0.658,0.404,0.764,0.596
---
BASE UNTUNED MODEL:
LFM2-8B-A1B
mxfp8 0.460,0.575,0.829,0.624,0.394,0.711,0.567