As with any Q3-30B-A3B, Designant performs very adequately with few or zero layers offloaded to GPU. When using the
ik_llama.cpp server, a 7950X CPU with 32GB of DDR5 RAM can run a Q4_K_M quant of this architecture at ~15 tokens/sec
with no GPU involved at all.