Gemma-4-E4B-IT is a ~4-billion-parameter instruction-tuned multimodal language model developed by Google, part of the Gemma 4 family. Its text backbone features
hybrid attention (36 of 42 layers sliding-window linear attention + 6 standard full-attention layers),
GeGLU activation,
Grouped Query Attention (GQA, 2 KV heads for 8 query heads),
per-layer input projections, and
128K native context length. Pre-trained on diverse web-scale corpora and aligned via instruction tuning + RLHF. This distribution is provided by
Aria Compute as an
aria-quant-bundle — a quantized package using
Hadamard pre-processing + uniform per-channel 8-bit quantization. Optimized for
CPU-only, on-device inference on mobile phones, edge devices, and single-board computers via the
Aria Engine runtime. No GPU or cloud connection is required.
Authenticated dashboard users can download the bundle via:
https://ariacompute.com/dashboard/models
Users (both direct and downstream) should be made aware of the above risks, biases, limitations, and constraints of the model. We recommend: