This issue appears to originate from GPTQModel, as it does not occur in the
E2B version. We are currently investigating and working on a fix.
This is an unofficial quantized version of google/gemma-4-31B-it.
FOEM is an improved quantization method over GPTQ. The resulting model preserves the same inference structure as GPTQ, ensuring compatibility with existing deployment pipelines while achieving better accuracy.