Gemma-4-E2B-IT is a ~1.5-billion-parameter instruction-tuned multimodal language model developed by Google, part of the Gemma 4 family. Its text backbone features
hybrid attention (27 of 35 layers sliding-window linear attention + 8 standard full-attention layers),
GeGLU activation,
aggressive Grouped Query Attention (GQA, 1 KV head for 8 query heads),
per-layer input projections,
double-wide MLP, and
128K native context length. Pre-trained on diverse web-scale corpora and aligned via instruction tuning + RLHF. This distribution is provided by
Aria Compute as an
aria-quant-bundle — a quantized package using
Hadamard pre-processing + uniform per-channel 8-bit quantization. Optimized for
CPU-only, on-device inference on mobile phones, edge devices, and single-board computers via the
Aria Engine runtime. No GPU or cloud connection is required.
Authenticated dashboard users can download the bundle via:
https://ariacompute.com/dashboard/models
Users (both direct and downstream) should be made aware of the above risks, biases, limitations, and constraints of the model. We recommend: