This repository provides a pruned version of the Gemma 4 E4B GGUF model, optimized for significantly lower RAM usage during inference.
The original Gemma 4 architecture introduces Per-Layer Embeddings, which improves performance but dramatically increases memory requirements—especially for multi-language support.
To address this, this version removes approximately ~45% of embedding weights corresponding to unused languages, resulting in: