After testing the new important matrix quants for 11b and 8x7b models and being able to run them on machines without a dedicated GPU, we are now exploring the middleground - 20b.
IQ3_S has been generated after PR
#5829 was merged. This should provide a significant speed boost even if you are offloading to CPU.
(Credits to
TeeZee for the original model and
ikawrakow for the stellar work on IQ quants)